Skip to main content

Samsara

Data Engineer

Remote - US (Remote)From $167kmidAdded today

About this role

Samsara seeks a Data Engineer to design and maintain robust data pipelines using SparkSQL and PySpark, transforming IoT and product data into analytics-ready datasets that power dashboards, model training, and statistical analysis across the Connected Operations Cloud platform.

What you'll do

  • Design and maintain reliable data pipelines using SparkSQL and PySpark within the central data lake
  • Build computed tables integrating data from multiple sources including unstructured video/audio, sensor data, and customer metadata
  • Deliver high-quality, customer-facing datasets with strong uptime and reliability standards
  • Collaborate with Data Science, Analytics, and AI/ML teams to ensure data quality for causal inference and model training
  • Access, manipulate, and integrate external datasets with internal data sources
  • Champion Samsara's cultural principles while embedding them in team practices at scale

What they're looking for

  • SparkSQL and PySpark
  • Data pipeline design and ETL development
  • SQL and Python
  • Data orchestration tools (Airflow, Dagster, or Prefect)
  • Large-scale data modeling
  • REST API integration
  • Git/GitHub version control
  • Software engineering fundamentals
Apply on the employer's site

Opens the official application on the employer’s site. No login required.

Samsara

Samsara builds an IoT operations platform that serves agriculture, construction, transportation, and manufacturing industries, with solutions for fleet management, safety, and efficiency across vehicle types. The company is hiring technical support engineers, sales engineers for public sector clients, solutions integration engineers, and software engineers to support its platform infrastructure and customer implementations.

View all jobs at Samsara

Likely interview questions

  • Walk us through a complex ETL pipeline you designed—what data sources did you integrate and how did you handle late-arriving or out-of-order data?
  • Describe your experience with Spark-based platforms. How have you optimized pipeline performance when dealing with large volumes of data?