Skip to main content

OpenAI

Operating Systems Engineer, On-Device Inference | Consumer Devices

San Francisco$230k–$385kfulltimemidAdded today

About this role

Design and develop the operating system stack that enables advanced AI inference on consumer devices with optimal performance, power efficiency, and responsiveness. You'll work across OS services, inference runtimes, model optimization, and resource scheduling to bring AI capabilities reliably into everyday products.

What you'll do

  • Design and implement OS services, frameworks, and interfaces for inference execution, model loading, and resource management
  • Optimize models for device constraints through quantization, runtime integration, and memory optimization in collaboration with researchers
  • Develop scheduling and resource policies that balance inference workloads with other device activities while maintaining latency and thermal limits
  • Create performance and power management strategies that adapt to workload needs and changing device conditions
  • Debug across the full stack using tracing, profiling, and structured analysis to identify correctness, concurrency, and reliability issues
  • Build diagnostic tools, instrumentation, and automated tests to validate performance gains and catch regressions on physical devices

What they're looking for

  • Operating system design and systems programming
  • C++ for concurrent programming and resource management
  • Inference runtime integration and optimization
  • Performance profiling and power analysis
  • Scheduling and memory management
  • Machine learning model optimization and quantization
  • Hardware accelerator optimization (CPUs, GPUs, neural processors)
  • Rust for systems programming (preferred)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Walk us through a time you optimized an inference workload in a resource-constrained environment—what were the key bottlenecks and how did you measure improvement?
  • Describe your experience integrating an inference runtime into an OS stack. What challenges arose around scheduling and resource contention?