Senior ML Inference Systems Engineer at Runpod
Confirmed open
Optimize LLM inference serving at Runpod across GPU hardware and workloads. Measure throughput, token latency and cost, profile bottlenecks from scheduling to kernels, and ship reliable production configurations. Remote work is listed in the United States.
- Location
- Remote
- Employment
- FullTime
- Work arrangement
- remote
- Technologies
- Python
Role overview
Optimize LLM inference serving at Runpod across GPU hardware and workloads. Measure throughput, token latency and cost, profile bottlenecks from scheduling to kernels, and ship reliable production configurations. Remote work is listed in the United States.
How to apply
Review the original posting and apply through the employer’s careers page.
View source and applyApply directlySource checked: 2026-10-02
Show what you can do
Build a free profile around your CV and real work. Share its link when you are ready. Applications for this role still go directly to the employer.
Build a free profile