AI Inference Platform Engineer
Quant trading · 36 postings open
- Location: Chicago
- Pay: Pay not stated
- Posted: Posted
About the role
The team builds and maintains the systems that serve AI models for the firm. The engineer optimizes inference performance, manages model onboarding, and ensures the reliability and efficiency of the serving platform across various hardware and software configurations.
Our summary of the posting; the original is on the DRW careers page.
What the posting requires
- Required
- NVIDIA GPU
- HBM
- TensorRT-LLM
- Continuous batching
- KV cache
- Nsight
- Linux
- Bottleneck diagnosis
- Hopper
- Tensor Cores
- Preferred
- GPU bottlenecks
- Serving decisions
- Evaluation harnesses
- Benchmarks
- Regression detection
These tags are our reading of the posting. They can miss something; the original posting is the reference.
Do you qualify?
Each requirement checked against your resume, in about a minute. No account. Your resume is deleted when the verdict appears.
