Save application

zettabyte-space

Oct 27

Senior/Staff Backend Engineer - Distributed System

Remote Remote · United States

pythondockerkubernetesgographql engineeringfulltime

ashby

Match & cover letter

Create a profile to see your match and get a cover letter for this role.

Create profile and match

Job description

Join our team to build systems that manage GPU clusters for AI workloads. You'll design APIs and develop resource management systems in a hybrid work environment.

Details

  • Hybrid role - 3 days in office, 2 days WFH; Must locate in Palo Alto
  • Competitive salary and equity based on experience and skillset
  • 5+ years backend engineering experience with distributed systems
  • Strong proficiency in Go, Python, or similar backend languages

The work

  • Design APIs that abstract complex GPU operations into simple developer experiences
  • Build scheduling algorithms that maximize GPU utilization while ensuring SLA compliance
  • Develop resource management systems for GPU lifecycle—provisioning, allocation, scheduling, and release
  • Create usage tracking and billing systems for GPU-hours, memory usage, and compute utilization