Samsara
Data Engineer
About this role
Samsara seeks a Data Engineer to design and maintain robust data pipelines using SparkSQL and PySpark, transforming IoT and product data into analytics-ready datasets that power dashboards, model training, and statistical analysis across the Connected Operations Cloud platform.
What you'll do
- Design and maintain reliable data pipelines using SparkSQL and PySpark within the central data lake
- Build computed tables integrating data from multiple sources including unstructured video/audio, sensor data, and customer metadata
- Deliver high-quality, customer-facing datasets with strong uptime and reliability standards
- Collaborate with Data Science, Analytics, and AI/ML teams to ensure data quality for causal inference and model training
- Access, manipulate, and integrate external datasets with internal data sources
- Champion Samsara's cultural principles while embedding them in team practices at scale
What they're looking for
- SparkSQL and PySpark
- Data pipeline design and ETL development
- SQL and Python
- Data orchestration tools (Airflow, Dagster, or Prefect)
- Large-scale data modeling
- REST API integration
- Git/GitHub version control
- Software engineering fundamentals
Opens the official application on the employer’s site. No login required.
Samsara
Samsara builds an IoT operations platform that serves agriculture, construction, transportation, and manufacturing industries, with solutions for fleet management, safety, and efficiency across vehicle types. The company is hiring technical support engineers, sales engineers for public sector clients, solutions integration engineers, and software engineers to support its platform infrastructure and customer implementations.
- Website
- samsara.com
Likely interview questions
- Walk us through a complex ETL pipeline you designed—what data sources did you integrate and how did you handle late-arriving or out-of-order data?
- Describe your experience with Spark-based platforms. How have you optimized pipeline performance when dealing with large volumes of data?