Benchling
Data Engineer
About this role
Benchling seeks a Data Engineer to build and operate production-grade data pipelines and warehouse infrastructure supporting the company's AI and analytics initiatives. You'll own end-to-end ELT workflows, manage Snowflake governance, and enable trustworthy data for internal AI applications across GTM, Product, Finance, and other departments.
What you'll do
- Build and operate production ELT pipelines ingesting data from Benchling's product, Salesforce, and third-party systems into Snowflake using dbt
- Ensure data quality, monitoring, testing, and schema versioning as usage scales
- Manage Snowflake access controls, data governance, PII handling, and warehouse cost/performance optimization
- Partner with AI engineering to make governed, trustworthy datasets available for agentic AI tooling
- Contribute to data architecture decisions including warehouse design and metrics store strategy
- Maintain pipeline health and support analytics needs across all business functions
What they're looking for
- SQL and Python
- dbt data modeling
- Snowflake or modern cloud data warehouse
- ELT pipeline design and orchestration (Airflow or similar)
- Data governance and RBAC
- Software engineering practices (version control, CI/CD, testing)
- Cloud infrastructure (AWS or equivalent)
- Data quality monitoring and schema management
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Benchling
Benchling builds an AI-powered platform for biotech R&D that integrates scientific workflows and data processes to accelerate research breakthroughs. The company is hiring software engineers across full-stack, customer engineering, agentic AI, and security roles to enhance developer productivity, build production AI systems, and protect sensitive research data.
- Website
- benchling.com
Likely interview questions
- Walk us through a production data pipeline you've built—what was the ingestion strategy, how did you model the data, and how do you ensure it stays reliable at scale?
- Describe your experience with dbt. How do you approach data modeling, testing, and managing schema changes in a production environment?