Dark Wolf Solutions
Data Engineer - Expert
About this role
Dark Wolf seeks an Expert Data Engineer to architect and lead enterprise-grade data platforms for defense and intelligence missions. You'll design petabyte-scale distributed systems, optimize processing frameworks, implement security controls, and mentor engineering teams in multi-cloud and air-gapped environments.
What you'll do
- Lead architectural design and implementation of distributed batch and streaming data platforms
- Optimize Spark, Trino/Presto, and other execution engines for petabyte-scale datasets under strict SLAs
- Provision data platform infrastructure using Infrastructure as Code across multi-cloud environments
- Establish zero-trust security, data lineage tracking, and automated compliance scanning for regulated data
- Act as Technical Pod Lead driving Agile deliverables, technical standards, and mentoring
What they're looking for
- Apache Spark, Flink, Kafka, Ray, and workflow orchestration
- Distributed systems internals, LSM trees, lakehouse formats (Delta Lake, Iceberg, Hudi)
- Python, Scala, SQL, and Go programming
- Kubernetes cluster management and custom operators
- Terraform and multi-cloud infrastructure architecture
- DevSecOps pipelines and automated compliance testing
- Enterprise data modeling and streaming schema contracts
- Technical leadership and Agile methodologies
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Dark Wolf Solutions
Dark Wolf Solutions builds DevSecOps platforms, integration systems, and data analytics solutions for defense and intelligence customers, with a focus on DoD and Space Force operations. The company is hiring DevOps engineers, full-stack software engineers, systems engineers, and data engineers to support cloud infrastructure, platform development, radar systems, and mission-critical defense applications.
View all jobs at Dark Wolf SolutionsLikely interview questions
- Describe your experience optimizing distributed query engines like Spark or Presto for petabyte-scale workloads—what specific bottlenecks did you address?
- How have you designed and implemented data mesh architectures, and what governance patterns did you establish?