Skip to main content

Garner Health

Data Engineer III

New York City, New York$166k–$205kfull-timemidAdded today

About this role

Garner seeks a Data Engineer III to build and optimize data pipelines supporting healthcare transformation across a 320M+ patient dataset. You'll design scalable data systems, create reusable datasets for BI and analytics teams, and establish data validation frameworks while ensuring privacy and security compliance.

What you'll do

  • Build, optimize, and maintain data pipelines powering business operations
  • Design and evangelize federated data validation framework for monitoring inconsistencies
  • Define and construct reusable, abstracted datasets for BI, Marketing, and Data Science
  • Ensure data security and HIPAA compliance across systems
  • Implement query optimization and data modeling best practices
  • Collaborate with cross-functional teams on data infrastructure scaling

What they're looking for

  • Python and SQL expertise
  • Apache Airflow orchestration
  • Snowflake data warehousing
  • PostgreSQL databases
  • Data pipeline design and optimization
  • HIPAA compliance and healthcare data security
  • Medallion/event-driven architectures
  • AWS cloud infrastructure

Benefits

  • Flexible PTO
  • Medical, Dental, and Vision plan options
  • 401(k) with company match
  • Flexible spending accounts
  • Teladoc Health access
  • Equity incentive participation
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Garner Health

Garner Health builds healthcare analytics systems that process medical data at scale using AI and distributed architectures. The company is hiring Software Engineers and Data Engineers to develop mission-critical infrastructure, data pipelines, and business intelligence solutions while maintaining privacy and security compliance.

View all jobs at Garner Health

Likely interview questions

  • Describe your experience building and optimizing data pipelines at scale—what frameworks and tools did you use?
  • How have you approached data quality monitoring and validation across distributed systems?