Skip to main content

Indicium AI

Data Engineer Consultant

New York$120k–$175kmidAdded today

About this role

Build and deploy production data pipelines for enterprise clients, modernizing legacy systems into cloud-native Lakehouses using Databricks, dbt, and Python. Work directly with clients and global engineering teams to deliver scalable data solutions in weeks.

What you'll do

  • Convert legacy data systems into high-performance PySpark and Databricks SQL pipelines using Medallion Architecture
  • Design and implement dbt models and SQL transformations to prepare clean, business-ready datasets
  • Provision and manage cloud infrastructure (AWS/GCP) using Terraform and enforce data governance with Unity Catalog
  • Build and monitor automated ELT/ETL workflows using Apache Airflow and Databricks Workflows with SLA compliance
  • Collaborate directly with clients in daily standups and code reviews, working with nearshore delivery teams
  • Debug failing pipelines, optimize complex SQL/Spark queries, and reduce cloud computing costs

What they're looking for

  • Python (production-grade code)
  • Advanced SQL
  • PySpark
  • Databricks (Delta Lake, Unity Catalog, Workflows)
  • dbt (data transformation, testing, documentation)
  • Terraform (Infrastructure as Code)
  • Apache Airflow
  • AWS or GCP cloud platforms
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Indicium AI

Indicium AI builds scalable data pipelines and infrastructure that support enterprise AI initiatives, including data integration, warehouse implementation, and ELT processes. The company is hiring Data Engineer Consultants to design and deliver production-grade data solutions while managing cloud infrastructure across cross-functional teams.

View all jobs at Indicium AI

Likely interview questions

  • Walk us through a legacy data system you've migrated to a modern Lakehouse architecture—what were the biggest challenges?
  • How do you approach designing dbt models for data quality and maintainability in large-scale projects?