← All jobs

Senior GPU Site Reliability Engineer at Andromeda

Confirmed open

Design and operate multi-provider GPU clusters for distributed AI training and inference. The role covers topology-aware scheduling, storage and networking, reliability targets, incident response and direct customer troubleshooting. The posting lists global remote work or San Francisco.

Andromeda initialsAndromeda
Location
Remote
Advertised region
Global Remote / San Francisco, CA / San Francisco / United States
Employment
FullTime
Work arrangement
remote
Technologies
Python, Kubernetes, PyTorch, Linux, Go

Role overview

Design and operate multi-provider GPU clusters for distributed AI training and inference. The role covers topology-aware scheduling, storage and networking, reliability targets, incident response and direct customer troubleshooting. The posting lists global remote work or San Francisco.

How to apply

Review the original posting and apply through the employer’s careers page.

View source and applyApply directly

Source checked: 2026-10-03

Show what you can do

Build a free profile around your CV and real work. Share its link when you are ready. Applications for this role still go directly to the employer.

Build a free profile

Browse job categories

Related jobs