H1
Data Engineer II- Life Sciences
New York$190k–$220kfull timemidAdded today
About this role
H1 seeks an experienced Data Engineer II to lead the design and optimization of distributed data pipelines for healthcare entity resolution on the EMERALD platform. You'll architect scalable Spark-based systems that link millions of healthcare records from diverse sources while mentoring a small engineering team and partnering with AI/ML and product teams to improve accuracy and operational efficiency.
What you'll do
- Lead design and optimization of Spark/PySpark pipelines for entity resolution and healthcare data processing at scale
- Own systems supporting automatching, identity mapping, deduplication, and enrichment workflows across healthcare provider datasets
- Build and maintain scalable frameworks for integrating PubMed, clinical trials, ct.gov, and other healthcare data sources
- Drive infrastructure optimization initiatives to improve throughput, cost efficiency, and observability
- Collaborate with AI/ML teams to integrate matching models and improve matching precision
- Mentor engineers through code reviews and technical guidance; support production operations and incident response
What they're looking for
- Apache Spark and PySpark
- Python, Scala, or Java for distributed systems
- AWS (EMR, S3, distributed compute)
- Entity resolution and identity matching systems
- ETL/ELT frameworks (batch and streaming)
- Kafka or Spark Streaming
- Docker, Kubernetes, Terraform, CI/CD pipelines
- Distributed systems design and performance optimization
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
H1
Likely interview questions
- Describe a complex entity resolution or deduplication challenge you've tackled in a previous role—what was your approach and how did you measure success?
- How have you optimized Spark job performance when processing tens of millions of records, and what metrics did you monitor?