Senior GPU Site Reliability Engineer at Andromeda
Confirmed open
Design and operate multi-provider GPU clusters for distributed AI training and inference. The role covers topology-aware scheduling, storage and networking, reliability targets, incident response and direct customer troubleshooting. The posting lists global remote work or San Francisco.
- Location
- Remote
- Advertised region
- Global Remote / San Francisco, CA / San Francisco / United States
- Employment
- FullTime
- Work arrangement
- remote
- Technologies
- Python, Kubernetes, PyTorch, Linux, Go
Role overview
Design and operate multi-provider GPU clusters for distributed AI training and inference. The role covers topology-aware scheduling, storage and networking, reliability targets, incident response and direct customer troubleshooting. The posting lists global remote work or San Francisco.
How to apply
Review the original posting and apply through the employer’s careers page.
View source and applyApply directlySource checked: 2026-10-03
Show what you can do
Build a free profile around your CV and real work. Share its link when you are ready. Applications for this role still go directly to the employer.
Build a free profile