Sarathi-Serve
OSDI'24Scheduler and execution co-design for better throughput and latency
I'm co-founder and CEO of Avartha. Today, real-time AI such as voice runs on small models because frontier models are too slow. We're rebuilding the inference stack so that real-time applications can use the most capable models.
Avartha spun out of the Systems for AI Lab at Georgia Tech, where I completed my CS PhD with Prof. Alexey Tumanov. Our scheduling and parallelism work, including chunked prefills from Sarathi-Serve, runs in vLLM, SGLang, and TensorRT-LLM. Before that, I built training infrastructure at Microsoft Research and Qubole.
Selected work across scheduling, capacity planning, long-context serving, and emulation.
Scheduler and execution co-design for better throughput and latency
Large-scale simulation for capacity planning and deployment exploration
Long-context serving across scheduling, parallelism, and memory
GPU-free time-warp emulation for reproducible serving experiments