Skip to main content

Baselayer

Data Engineer

San Francisco, California$120k–$150kmidAdded today

About this role

Baselayer is seeking a Data Engineer to build and maintain ETL/ELT pipelines that power a real-time business identity graph used by major financial institutions. You'll work with public records, fraud telemetry, and web signals to create production-grade data infrastructure that enables entity resolution and fraud detection at scale.

What you'll do

  • Build and maintain ETL/ELT pipelines ingesting data from dozens of sources including public records and fraud telemetry
  • Develop data models and transformation layers using Dataflow, Spark, and Airflow for fraud detection and KYB APIs
  • Implement data quality checks, observability, and alerting to catch issues before customer impact
  • Optimize pipelines for performance, freshness, and cost in cloud data warehouses
  • Collaborate with data scientists and ML engineers to provide clean, well-modeled data for entity resolution
  • Ensure pipelines meet security and regulatory standards for sensitive data (SOC 2, GDPR, KYC/KYB)

What they're looking for

  • Python
  • SQL
  • ETL/ELT pipeline development
  • Dataflow, Spark, and/or Airflow
  • Cloud data warehouses (BigQuery, Snowflake)
  • Data modeling and integrity
  • Streaming systems (Kafka, Pub/Sub)
  • GCP cloud platform
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Baselayer

Baselayer builds identity verification solutions that help customers integrate and implement technical systems. The company is hiring Solutions Engineers to bridge engineering and customer success by facilitating technical integrations and translating complex concepts for diverse audiences.

View all jobs at Baselayer

Likely interview questions

  • Walk us through a complex ETL pipeline you built—what data sources did you integrate and how did you handle quality and freshness?
  • Describe your approach to debugging a production pipeline failure affecting downstream models or APIs.