Data Scientist II, ML Infrastructure
About this role
Pinterest is seeking a Data Scientist II to advance ML measurement, feature understanding, and causal inference at scale. You'll translate research into production ML systems, build self-serve tooling for causal analysis, and create centralized platform capabilities that improve ML infrastructure across the organization.
What you'll do
- Convert research workflows into production ML pipelines using Airflow, WandB, and Ray while establishing reusable patterns
- Productionize causal inference methods (propensity scoring, IPW, TMLE) and build self-serve tooling for non-experts
- Partner with ML engineers and product teams to identify tooling and measurement improvements
- Leverage platform metadata to build data-driven frameworks for feature importance and content deindexing
- Design and operate centralized ML platform tooling for feature creation, model evaluation, and trust at scale
What they're looking for
- Python and PyTorch or equivalent deep learning frameworks
- Distributed computing (Spark, Ray)
- Causal inference methodologies
- Workflow orchestration (Airflow, Prefect, Jenkins)
- Software development best practices and version control
- ML theory and first-principles reasoning
- Production ML systems design
Opens the official application on the employer’s site. No login required.
Pinterest builds a large-scale platform serving millions of users, with infrastructure spanning security, database systems, and mobile products, supported by data systems and advertising technology. The company is hiring Software Engineers II and experienced engineers across security, infrastructure, iOS development, and data engineering to enhance platform capabilities, improve detection and response systems, optimize performance, and leverage AI-driven solutions.
- Website
- pinterest.com
Likely interview questions
- Describe a research project you've taken from prototyping to production—what were the key challenges in operationalizing it?
- How would you approach building a self-serve platform for causal inference that non-ML experts could use reliably?