Skip to main content

OpenAI

Data Engineer, People Innovation Labs

San Francisco (Remote)$293k–$325kfulltimemidAdded 1 month ago

About this role

OpenAI's People Innovation Labs seeks a Data Engineer to design and maintain data pipelines supporting internal HR products and people analytics. You'll build scalable data systems that integrate employee information into the Databricks warehouse and enable data-driven insights across recruiting, culture, and organizational initiatives.

What you'll do

  • Design and manage people data pipelines integrating with Databricks warehouse
  • Develop canonical datasets tracking people metrics and product performance
  • Collaborate with Data Platform, Analytics, and People teams to understand data requirements
  • Build robust, fault-tolerant systems for data ingestion and processing
  • Lead data architecture decisions as the primary engineering expert on the team
  • Ensure data security, integrity, and compliance with industry standards

What they're looking for

  • Data engineering (3+ years experience)
  • Python, Scala, or Java programming
  • Databricks and/or Snowflake
  • ETL schedulers (Airflow, Fivetran, Dagster, Prefect, or similar)
  • Distributed processing (Spark, Hadoop, Flink)
  • Distributed storage systems (HDFS, S3)
  • Data warehouse design
  • Software engineering fundamentals (8+ years)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Walk us through your experience designing and building data pipelines. What technologies did you use, and how did you ensure fault tolerance and reliability?
  • Tell us about your experience with Databricks or similar data warehousing platforms. How have you optimized data ingestion and processing at scale?