Save application

Kimchi

1 day ago

Senior ML Engineer | Kimchi (LLM Inference Optimization)

Onsite

pythonnodeawskubernetesgcpazurepytorch technology

arbeitnow

Match & cover letter

Create a profile to see your match and get a cover letter for this role.

Create profile and match

Job description

You will optimize LLM inference performance at Kimchi, focusing on throughput and latency. This role allows you to lead technical decisions and work on cutting-edge cloud-native infrastructure.

Details

  • 5+ years of experience in ML systems
  • Location: Miami, in-office
  • Skills: Python, vLLM, SGLang, TensorRT-LLM, Kubernetes, AWS, GCP, Azure

The work

  • Push throughput through kernel-level tuning and continuous batching
  • Cut latency by profiling and fixing bottlenecks
  • Optimize KV cache utilization with advanced techniques
  • Quantize models without quality regression