Skip to content

Software Engineer- Inference Performance

Baseten

AI · 57 roles open

  • Location: San Francisco
  • Pay: $180k–$360k a year
  • Posted: Posted
Check your fit

About the role

The team optimizes AI inference performance across the entire stack, from runtime to serving. The engineer analyzes and improves the speed and efficiency of demanding AI workloads, directly impacting customer model performance and resource utilization.

Our summary of the posting; the original is on the Baseten careers page.

What the posting requires

Required
  • Python
  • LLM optimization
  • ML libraries
  • GPU architecture
  • C++
  • Quantization
  • PyTorch
  • Programming languages
  • Speculative decoding
  • TensorRT
Preferred
  • Software performance
  • vLLM
  • GPU kernels
  • Software engineering principles
  • LLM performance
  • SGLang

These tags are our reading of the posting. They can miss something; the original posting is the reference.

Do you qualify?

Each requirement checked against your resume, in about a minute. No account. Your resume is deleted when the verdict appears.

Check your fit