OpenAI
Data Engineer, Monetization Data Platform
About this role
Build and operate large-scale data pipelines and platforms that power monetization decisions across product, finance, and GTM teams. Own end-to-end systems from instrumentation through modeling, quality controls, and delivery to downstream consumers, working across complex financial and operational data domains.
What you'll do
- Design and operate streaming and batch data pipelines processing product, financial, and operational data from multiple sources
- Develop canonical data models and reusable products for pricing, billing, payments, revenue, and general ledger domains
- Establish data accuracy, completeness, freshness, lineage, and auditability guarantees
- Build platform frameworks and capabilities to improve developer productivity for monetization features
- Partner with Product, Finance, Accounting, and GTM teams to define data contracts and instrument new monetization products
- Lead technical design of complex cross-functional projects and improve observability of critical data workflows
What they're looking for
- Data pipeline architecture and design
- Python, Java, or Scala programming
- Data modeling and distributed systems
- Data quality, governance, and observability
- SQL and data transformation frameworks
- Workflow orchestration and scheduling
- Streaming systems (Kafka, Pub/Sub, etc.)
- Cross-functional collaboration and communication
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Describe a complex data pipeline you've built end-to-end—what were the key challenges around data quality and lineage?
- How do you approach designing data models when you have conflicting requirements from multiple downstream teams?