OpenAI
Research Engineer/Research Scientist - Personal AGI, North Stars
About this role
Join OpenAI's Personal AGI team to advance AI capabilities that benefit millions of users globally. You'll conduct research on model behavior, safety, and personalization, designing evaluations and improvements to close gaps between power users and everyday consumers through better tool-use, instruction following, and intelligent features.
What you'll do
- Own and execute research initiatives to enhance model capability and performance
- Build and maintain robust evaluation frameworks for tracking model improvements
- Design, implement, and debug code across the research infrastructure stack
- Collaborate with research and product teams to optimize customer experiences
- Analyze model bottlenecks and translate findings into training data and reward signals
- Contribute to safety, factuality, and personalization improvements
What they're looking for
- Machine learning engineering and research
- Large language model knowledge and applications
- Evaluation design and methodology
- Python and ML codebase debugging
- Data and reward signal engineering
- Product-driven research mindset
- Cross-functional collaboration
- Technical problem-solving in complex environments
Benefits
- Hybrid work model (3 days in-office per week)
- San Francisco-based role
- Relocation assistance provided
- Work on frontier AI technology
- Collaborate with leading AI research teams
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Can you walk us through a time when you identified a bottleneck in model behavior and how you translated that insight into evals, training data, or reward signals?
- Describe your experience building and validating evaluations for measuring model capabilities. What metrics have you found most useful?