Save application

build-ai

Aug 30

ML Engineer, Inference Optimization

Onsite San Francisco

pythonrustpytorch softwarefulltime

ashby

Match & cover letter

Create a profile to see your match and get a cover letter for this role.

Create profile and match

Job description

You will optimize inference performance to reduce costs and improve efficiency in San Francisco. This role involves working closely with research and product teams to ensure scalable model deployment.

Details

  • San Francisco, work from office
  • Competitive pay
  • Experience in ML/systems engineering with inference optimization
  • Skills in Python and C++ or Rust

The work

  • Own inference performance metrics like latency and cost
  • Optimize compute through techniques like batching and quantization
  • Profile pipelines to identify and fix bottlenecks
  • Collaborate with teams to ensure models are cost-effective at scale