Capital Technology Group
Junior Data Engineer
About this role
Capital Technology Group seeks a Junior Data Engineer to design and maintain scalable data pipelines and ETL workflows for federal government clients. You'll work with modern cloud technologies including Databricks, AWS services, and Apache Spark to deliver mission-critical data solutions in a remote, collaborative agile environment.
What you'll do
- Design, build, and maintain scalable data pipelines and ETL/ELT workflows using Databricks, dbt, PySpark, SQL, and Python
- Develop AWS-native data solutions leveraging Glue, EMR, MWAA, Lambda, S3, Redshift, and related services
- Build high-performance data ingestion and transformation workflows across structured and semi-structured data formats
- Integrate data from enterprise and external sources including relational and NoSQL databases
- Monitor, troubleshoot, and optimize enterprise data platforms for reliability, scalability, and performance
- Collaborate with cross-functional teams in Agile sprints to define requirements and deliver data solutions
What they're looking for
- Apache Spark and PySpark
- SQL and database design (PostgreSQL, Oracle, Redshift)
- Databricks and dbt
- AWS services (Glue, EMR, MWAA, Lambda, Step Functions, S3)
- Data pipeline orchestration and workflow automation
- Python programming
- Cloud platform architecture
- Agile methodologies
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Capital Technology Group
Capital Technology Group builds secure, scalable technology solutions for federal government modernization initiatives, with particular expertise in Identity & Access Management systems. The company is hiring quality assurance and solutions architecture roles to lead technical delivery on mission-critical government programs.
View all jobs at Capital Technology GroupLikely interview questions
- Walk us through how you would design a scalable data pipeline to ingest data from multiple enterprise sources into Redshift, including your approach to handling data quality and transformation.
- Describe your experience with Apache Spark or PySpark—what performance optimization techniques have you used when working with large datasets?