Skip to content

Engineering Manager - Inference Performance

Baseten

AI · 57 roles open

  • Location: San Francisco
  • Pay: $240k–$270k a year
  • Posted: Posted
Check your fit

About the role

The Inference Performance team enhances the speed and efficiency of complex AI computations on GPUs. The Engineering Manager guides and expands a group working on the inference engine and runtime. This role includes defining technical strategy, resolving difficult challenges, and supporting the team that boosts customer model execution.

Our summary of the posting; the original is on the Baseten careers page.

What the posting requires

Required
  • GPU architecture
  • PyTorch
  • Performance tradeoffs
  • TensorRT
  • TensorRT-LLM
  • ML libraries
  • GPU workloads
  • Bachelor
  • Master
  • PhD
Preferred
  • vLLM
  • GPU kernels
  • SGLang
  • CUDA
  • Triton
  • Inference engines

These tags are our reading of the posting. They can miss something; the original posting is the reference.

Do you qualify?

Each requirement checked against your resume, in about a minute. No account. Your resume is deleted when the verdict appears.

Check your fit