Save application

castaigroupinc

11 days ago

Senior ML Engineer | Kimchi (LLM Inference Optimization)

Onsite Austria; France; Germany; Italy; Netherlands; Poland; Spain; United Kingdom

pythonnodeawskubernetesgcpazurepytorch technology

greenhouse

Match & cover letter

Create a profile to see your match and get a cover letter for this role.

Create profile and match

Job description

You will optimize LLM inference performance for customers while leading technical direction at Cast AI. This role requires deep understanding of both machine learning systems and infrastructure.

Details

  • Location: Austria, France, Germany, Italy, Netherlands, Poland, Spain, United Kingdom
  • 5+ years of experience in ML systems
  • Skills: Python, vLLM, SGLang, TensorRT-LLM, Kubernetes, AWS, GCP, Azure

The work

  • Push throughput through kernel-level tuning and continuous batching
  • Cut latency by profiling and fixing bottlenecks
  • Optimize KV cache utilization for better throughput
  • Quantize models without quality regression