Software Engineer- Inference Performance
AI · 57 roles open
- Location: San Francisco
- Pay: $180k–$360k a year
- Posted: Posted
About the role
The team optimizes AI inference performance across the entire stack, from runtime to serving. The engineer analyzes and improves the speed and efficiency of demanding AI workloads, directly impacting customer model performance and resource utilization.
Our summary of the posting; the original is on the Baseten careers page.
What the posting requires
- Required
- Python
- LLM optimization
- ML libraries
- GPU architecture
- C++
- Quantization
- PyTorch
- Programming languages
- Speculative decoding
- TensorRT
- Preferred
- Software performance
- vLLM
- GPU kernels
- Software engineering principles
- LLM performance
- SGLang
These tags are our reading of the posting. They can miss something; the original posting is the reference.
Do you qualify?
Each requirement checked against your resume, in about a minute. No account. Your resume is deleted when the verdict appears.
