Save application

gimlet

Mar 10

Member of Technical Staff - ML Systems & Inference

Onsite San Francisco, CA

python research and developmentfulltime

ashby

Match & cover letter

Create a profile to see your match and get a cover letter for this role.

Create profile and match

Job description

You will build and optimize ML inference systems for production at Gimlet in San Francisco. This role involves working on batching, scheduling, and resource utilization for AI workloads.

Details

  • San Francisco, work from office
  • Experience building or operating ML inference or model serving systems
  • Bachelor's degree in a relevant field or equivalent experience
  • Skills in Python and C++

The work

  • Build inference systems that execute models end-to-end in production
  • Design execution strategies for batching, scheduling, and resource utilization
  • Improve KV cache management and memory efficiency
  • Enable new models and inference techniques to run efficiently