Clera
Research Engineer, Privacy and Anonymization
About this role
Build privacy-preserving data infrastructure that detects and removes sensitive information from raw data before it enters AI training and processing workflows. You'll own detection systems combining rules, ML classifiers, and LLMs, create evaluation frameworks measuring privacy risk versus data utility, and ensure robustness across diverse data sources and edge cases.
What you'll do
- Design and implement systems to detect PII, quasi-identifiers, credentials, and sensitive information using rules, statistical models, classifiers, and LLM-based approaches
- Develop production anonymization pipelines that transform data before downstream processing, training, evaluation, and synthetic data generation
- Build evaluation frameworks measuring privacy risk, data utility retention, leakage, and adversarial re-identification resilience
- Engineer systems resilient to schema drift, unusual formats, and sensitive data in unexpected fields
- Benchmark and compare detection approaches across recall, precision, latency, cost, and downstream utility metrics
- Collaborate with engineering, research, operations, and customers to translate privacy requirements into technical policies
What they're looking for
- Python production data and ML systems
- Information extraction and named-entity recognition
- Data classification and sensitive content detection
- End-to-end data pipeline architecture
- Privacy techniques (redaction, masking, pseudonymization, anonymization, synthetic data)
- Privacy-enhancing technologies (differential privacy, k-anonymity, format-preserving encryption)
- Low-latency ML inference and high-throughput data processing
- Experimental design and metrics evaluation
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Clera
Clera builds an agentic operating system that automates complex workflows and processes through AI agents, with a platform designed to simplify distributed infrastructure management for developers. The company is hiring Founding Engineers, Customer Engineers, and Product Engineers to develop both backend systems and user-facing interfaces across their AI automation products.
View all jobs at CleraLikely interview questions
- Describe a production data pipeline you built end-to-end—what were the biggest challenges and how did you handle schema changes?
- How would you design a system to detect both obvious PII and subtle quasi-identifiers that could enable re-identification?