Capital Technology Group
Data Engineer
About this role
Capital Technology Group seeks a Data Engineer to design and maintain scalable data pipelines and analytics platforms for federal clients. You'll work with modern tools like Apache Spark, Databricks, AWS services, and AI technologies to solve mission-critical data challenges in a remote, collaborative environment.
What you'll do
- Design and build scalable data pipelines, ETL/ELT workflows, and data models using Python, PySpark, Databricks, dbt, and SQL
- Develop and optimize AWS-native data platforms leveraging Glue, EMR, MWAA, Lambda, Step Functions, S3, Redshift, and RDS
- Build high-performance ingestion and transformation workflows for structured and semi-structured data using Apache Iceberg, Parquet, and Avro
- Design analytical platforms using Amazon Athena, Trino, Hive, and OpenSearch with enterprise data catalog technologies
- Integrate enterprise and external data sources across relational and NoSQL databases including PostgreSQL, Oracle, and GraphDB
- Build AI-enabled data solutions using Amazon Bedrock, RAG pipelines, and vector search technologies
What they're looking for
- Python and PySpark
- Apache Spark and Databricks
- AWS services (Glue, EMR, MWAA, Lambda, S3, Redshift, RDS)
- SQL and relational databases (PostgreSQL, Oracle)
- Data modeling and dbt
- CloudFormation and Infrastructure as Code
- Data formats (Apache Iceberg, Parquet, Avro, ORC)
- AI/ML integration (Amazon Bedrock, RAG, vector search)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Capital Technology Group
Capital Technology Group builds secure, scalable technology solutions for federal government modernization initiatives, with particular expertise in Identity & Access Management systems. The company is hiring quality assurance and solutions architecture roles to lead technical delivery on mission-critical government programs.
View all jobs at Capital Technology GroupLikely interview questions
- Describe your experience designing and optimizing ETL pipelines at scale—what challenges did you encounter and how did you solve them?
- How have you used Apache Spark or PySpark to handle large datasets, and what performance optimization techniques did you apply?