Anthropic
Cyber Evaluations Engineer
About this role
Anthropic is seeking a Cyber Evaluations Engineer to design and execute evaluations that assess cyber-related risks and safeguards in AI models. You'll analyze jailbreaks and prompt bypasses, build detection systems for cyber misuse, and collaborate with policy teams to strengthen model robustness across pre-release testing and production.
What you'll do
- Design and run capability, uplift, and safety evaluations for cyber-relevant risks in new models
- Execute per-release safeguard-robustness testing ahead of major model launches
- Analyze evaluation results and communicate findings to team and stakeholders
- Design and tune detection probes for identifying cyber misuse
- Build layered abuse-detection architecture in collaboration with policy teams
- Maintain internal tooling for running and scoring evaluations
What they're looking for
- Evaluation and benchmark design for software or ML systems
- Hands-on cybersecurity experience (CTF, vulnerability research, exploit development)
- Python proficiency
- Adversarial data analysis (jailbreaks, prompt bypasses, intrusion telemetry)
- Cross-functional stakeholder communication
- AI/ML evaluation frameworks
- Detection rule authoring (Sigma, YARA, Suricata, SIEM rules)
- Security clearance eligibility
Benefits
- Work on mission-critical AI safety and security challenges
- Collaborative environment with security researchers and policy experts
- Flexible hybrid work arrangement (minimum 25% office time)
- Visa sponsorship available
- Competitive annual compensation
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Anthropic
Anthropic builds Claude, an AI assistant, and is hiring for engineering roles across infrastructure, data systems, and security that support both AI research operations and the company's internal technology needs. The company seeks infrastructure engineers, systems integrators, data scientists, and security specialists to build production-scale systems for training data pipelines, financial operations, developer productivity measurement, research infrastructure, and server firmware security.
- Website
- anthropic.com
Likely interview questions
- Describe a significant vulnerability or security issue you've discovered or researched—what was your approach and how did you validate it?
- Walk us through how you'd design an evaluation to test whether an AI model could be jailbroken into providing harmful cyber advice.