Skip to content

Senior Staff Applied AI Inference Engineer

Crusoe

AI · 25 roles open in software and data

  • Location: San Francisco, CA - US
  • Pay: Pay not stated
  • Posted: Posted
Check your fit

About the role

The team works on the inference stack for large language models, aiming to make them faster and more cost-effective. The engineer optimizes serving architectures, profiles code, and collaborates with customer teams to deploy and monitor these models in production environments.

Our summary of the posting; the original is on the Crusoe careers page.

What the posting requires

Required
  • LLM optimization
  • LLM serving frameworks
  • Profiling
  • Python
  • High throughput
  • vLLM
  • Performance analysis
  • Low latency
  • SGLang
  • Kernel level
Preferred
  • CUDA
  • Docker
  • Kubernetes
  • Software performance
  • AI/ML inference systems
  • AI/ML projects

These tags are our reading of the posting. They can miss something; the original posting is the reference.

Do you qualify?

Each requirement checked against your resume, in about a minute. No account. Your resume is deleted when the verdict appears.

Check your fit