Skip to main content

OpenAI

ChatGPT Performance Engineer

San Francisco$325k–$405kfulltimemidAdded 1 month ago

About this role

OpenAI seeks an experienced Performance Engineer to optimize the infrastructure and application performance of ChatGPT and the OpenAI API at scale. You'll conduct root-cause analysis and drive systemic improvements across networking, storage, runtime, and GPU layers while collaborating with cross-functional teams.

What you'll do

  • Analyze and optimize performance across application, middleware, runtime, and infrastructure layers
  • Develop tooling and metrics for deep system performance observability
  • Collaborate with infrastructure, platform, training, and product teams on performance goals
  • Influence architecture decisions to prioritize latency, throughput, and efficiency
  • Lead investigations into production performance regressions and scalability issues
  • Define performance testing strategies and SLAs/SLOs for critical systems

What they're looking for

  • Performance profiling and tracing systems
  • Distributed systems optimization at scale
  • OS internals and memory management
  • Database, networking, and storage optimization
  • Python runtime and GPU utilization
  • Observability and benchmarking infrastructure
  • Root-cause analysis and debugging
  • Systems design and architecture
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Walk us through a time you identified and resolved a critical performance bottleneck in a distributed system. What tools and methodology did you use?
  • Describe your experience with performance profiling tools and tracing systems. Which ones have you used most deeply, and what insights did they reveal?