Software Engineer, Inference – AMD GPU Enablement
OpenAI · San Francisco, California, US
About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprises a...
Job description
About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprises and developers alike to use and access our state-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role: We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across emerging GPU platforms. You’ll work across the stack - from low-level kernel performance to high-level distributed execution - and collaborate closely with research, infra, and performance teams to ensure our largest models run smoothly on new hardware. This is a high-impact opportunity to shape OpenAI’s multi-platform inference capabilities from the ground up with a particular focus on advancing inference performance on AMD accelerators. In this role, you will: Own bring-up, correctness and performance of the OpenAI inference stack on AMD hardware. Integrate internal model-serving infrastructure (e.g., vLLM, Triton) into a variety of GPU-backed systems. Debug and optimize di...