You will optimize LLM inference performance for customers while leading technical direction at Cast AI. This role requires deep understanding of both machine learning systems and infrastructure.
Details
-
Location: Austria, France, Germany, Italy, Netherlands, Poland, Spain, United Kingdom
-
5+ years of experience in ML systems
-
Skills: Python, vLLM, SGLang, TensorRT-LLM, Kubernetes, AWS, GCP, Azure
The work
-
Push throughput through kernel-level tuning and continuous batching
-
Cut latency by profiling and fixing bottlenecks
-
Optimize KV cache utilization for better throughput
-
Quantize models without quality regression