You will support ML engineers with large-scale training and inference workloads in production environments. This role involves diagnosing complex issues and improving platform reliability.
Details
-
New York, San Francisco, or Seattle, in-office 2 days per week
-
$115,000—$140,000 USD
-
Experience with Kubernetes, cloud infrastructure, and distributed systems
-
No visa sponsorship available
The work
-
Partner with customer engineering teams to resolve ML infrastructure issues
-
Investigate failures in distributed training and Kubernetes orchestration
-
Analyze logs and system behavior to isolate root causes
-
Identify patterns in customer issues to drive reliability improvements