Senior Staff Applied AI Inference Engineer
AI · 25 roles open in software and data
- Location: San Francisco, CA - US
- Pay: Pay not stated
- Posted: Posted
About the role
The team works on the inference stack for large language models, aiming to make them faster and more cost-effective. The engineer optimizes serving architectures, profiles code, and collaborates with customer teams to deploy and monitor these models in production environments.
Our summary of the posting; the original is on the Crusoe careers page.
What the posting requires
- Required
- LLM optimization
- LLM serving frameworks
- Profiling
- Python
- High throughput
- vLLM
- Performance analysis
- Low latency
- SGLang
- Kernel level
- Preferred
- CUDA
- Docker
- Kubernetes
- Software performance
- AI/ML inference systems
- AI/ML projects
These tags are our reading of the posting. They can miss something; the original posting is the reference.
Do you qualify?
Each requirement checked against your resume, in about a minute. No account. Your resume is deleted when the verdict appears.
