Engineering Manager - Inference Performance
AI · 57 roles open
- Location: San Francisco
- Pay: $240k–$270k a year
- Posted: Posted
About the role
The Inference Performance team enhances the speed and efficiency of complex AI computations on GPUs. The Engineering Manager guides and expands a group working on the inference engine and runtime. This role includes defining technical strategy, resolving difficult challenges, and supporting the team that boosts customer model execution.
Our summary of the posting; the original is on the Baseten careers page.
What the posting requires
- Required
- GPU architecture
- PyTorch
- Performance tradeoffs
- TensorRT
- TensorRT-LLM
- ML libraries
- GPU workloads
- Bachelor
- Master
- PhD
- Preferred
- vLLM
- GPU kernels
- SGLang
- CUDA
- Triton
- Inference engines
These tags are our reading of the posting. They can miss something; the original posting is the reference.
Do you qualify?
Each requirement checked against your resume, in about a minute. No account. Your resume is deleted when the verdict appears.
