OpenAI
Software Engineer, Trainium
About this role
OpenAI seeks a Software Engineer to optimize AI model inference on AWS Trainium, working across kernels, compilers, and runtime systems. You'll develop high-performance software to efficiently execute frontier models on specialized hardware, collaborating with cross-functional teams to unlock accelerator capabilities.
What you'll do
- Build and optimize OpenAI's inference stack specifically for AWS Trainium hardware
- Develop high-performance kernels for critical model operations and workloads
- Extend compiler support to efficiently target Trainium, improving code generation
- Profile workloads and identify performance bottlenecks across the software stack
- Partner with inference and ML systems teams to integrate new model architectures
- Own end-to-end performance and systems problems from investigation through production deployment
What they're looking for
- Systems programming and performance-critical software development
- ML systems, compilers, or kernel development experience
- Specialized accelerator architecture knowledge (GPU, TPU, Trainium, or similar)
- Performance profiling and optimization across multiple abstraction layers
- Compiler infrastructure (LLVM, MLIR, XLA, or Triton) or ML framework experience
- AWS Trainium or AWS Neuron SDK familiarity (bonus)
- Hardware/software co-optimization and cross-boundary problem solving
- Experience with PyTorch, JAX, or kernel development for accelerators (bonus)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Describe your experience optimizing software for specialized accelerators like GPUs or TPUs—what were the key performance bottlenecks you tackled?
- Walk us through a time you owned a complex performance problem end-to-end. How did you approach profiling and identify the root cause?