OpenAI
Research Engineer / Research Scientist- Personal AGI (Post Training)
About this role
Join OpenAI's Post-Training team to research and develop improvements for large language models deployed in ChatGPT and the API. You'll combine reinforcement learning with product-driven research to enhance model capability, safety, and reliability for millions of users.
What you'll do
- Own and pursue a research agenda to improve model capability and performance
- Collaborate with research and product teams to optimize model deployment
- Build robust evaluation frameworks for tracking modeling improvements
- Design, implement, test, and debug code across the research stack
- Prepare models for real-world deployment ensuring safety and efficiency
What they're looking for
- Machine learning fundamentals and applications
- Reinforcement learning
- ML model evaluation and benchmarking
- Large codebase debugging and navigation
- ML engineering and implementation
- Python or similar programming languages
- Knowledge of modern language models
Benefits
- Hybrid work model (3 days in office per week)
- Based in San Francisco, CA
- Relocation assistance for new employees
- Work on cutting-edge AI systems
- Collaboration with leading AI research teams
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Describe your experience with reinforcement learning. How have you applied RL to improve model performance in a production or deployed setting?
- Walk us through how you've designed and built evaluation frameworks to measure model improvements. What metrics have you found most useful?