OpenAI
Software Engineer, Inference - Performance Optimization
About this role
OpenAI seeks a Software Engineer to optimize inference performance by building performance models, analyzing workloads across application and infrastructure layers, and creating tools to help teams understand latency, capacity, and cost tradeoffs. You'll partner with cross-functional teams to turn performance insights into concrete improvements for production systems.
What you'll do
- Develop and refine performance models converting microbenchmarks into cost-to-serve estimates
- Conduct end-to-end inference workload analysis across applications, models, and fleet infrastructure
- Build tooling to identify performance bottlenecks across abstraction layers for latency and throughput
- Partner with engineering and research teams to implement performance improvements
- Project how future system changes affect inference performance and capacity needs
- Analyze inference stack performance across application, model, and fleet layers
What they're looking for
- Performance profiling and benchmarking
- Distributed systems design and analysis
- Model inference optimization
- Hardware efficiency understanding
- Cross-layer system analysis (kernels, accelerators, networking)
- Systems modeling and capacity planning
- Data analysis and visualization
- Collaboration with engineering teams
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Walk us through how you would build a performance model that translates microbenchmark results into end-to-end cost-to-serve estimates across application, model, and fleet layers.
- Describe your experience with systems profiling and benchmarking tools. What challenges have you faced in identifying bottlenecks across different abstraction layers?