Truveta
Software Engineer - Data Processing
About this role
Join Truveta's data engineering team to design and build large-scale distributed data processing pipelines for healthcare applications. You'll work on ETL systems, cloud infrastructure, and solve complex distributed computing challenges while collaborating with product, clinical, and engineering teams in our Seattle headquarters.
What you'll do
- Design and implement large-scale distributed data processing pipelines and ETL solutions
- Contribute across the full software development lifecycle including design, testing, deployment, and performance optimization
- Build and maintain data pipeline systems capable of processing healthcare data at unprecedented scale
- Collaborate with product managers, clinical informaticists, architects, and cross-functional engineering teams
- Debug and resolve complex production issues in distributed systems
- Develop production-quality, multi-threaded code for cloud environments
What they're looking for
- Distributed systems design and implementation
- Cloud platforms (Azure preferred, AWS/GCP relevant)
- Data pipeline and ETL development
- Multi-threaded programming and concurrency
- DevOps practices and cloud-native architectures
- Large-scale data storage and distribution
- Production debugging and troubleshooting
- Software development lifecycle best practices
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Truveta
Truveta builds a healthcare data platform that leverages software engineering and AI to transform how healthcare organizations manage and derive insights from data. The company is hiring software engineers across multiple specializations including person matching, core platform services, and backend systems, as well as machine learning engineers focused on generative AI and large language models.
- Website
- truveta.com
Likely interview questions
- Describe a time you designed and implemented a large-scale distributed system. What challenges did you face and how did you overcome them?
- How have you optimized data pipeline performance when processing massive datasets?