Skip to content

AI Inference Platform Engineer

DRW

Quant trading · 36 postings open

  • Location: Chicago
  • Pay: Pay not stated
  • Posted: Posted
Check your fit

About the role

The team builds and maintains the systems that serve AI models for the firm. The engineer optimizes inference performance, manages model onboarding, and ensures the reliability and efficiency of the serving platform across various hardware and software configurations.

Our summary of the posting; the original is on the DRW careers page.

What the posting requires

Required
  • NVIDIA GPU
  • HBM
  • TensorRT-LLM
  • Continuous batching
  • KV cache
  • Nsight
  • Linux
  • Bottleneck diagnosis
  • Hopper
  • Tensor Cores
Preferred
  • GPU bottlenecks
  • Serving decisions
  • Evaluation harnesses
  • Benchmarks
  • Regression detection

These tags are our reading of the posting. They can miss something; the original posting is the reference.

Do you qualify?

Each requirement checked against your resume, in about a minute. No account. Your resume is deleted when the verdict appears.

Check your fit