SpyCloud
Data Processing Engineer
Austin, Texas | Remote (Remote)$111k–$144kfull-timemidAdded yesterday
About this role
SpyCloud seeks a Security Data Analyst/Python Developer to parse, transform, and automate processing of large datasets for cybersecurity applications. This hybrid/remote role requires 5-7 years of Python experience, AWS proficiency, and expertise in data cleaning and ETL pipelines to support the company's mission of disrupting cybercrime.
What you'll do
- Parse, clean, and transform structured and unstructured datasets to align with SpyCloud schema
- Build Python-based automation for the data parsing platform and develop ETL scripts
- Monitor data ingestion processes and troubleshoot issues
- Develop CI/CD pipelines and establish standards for codebase management
- Collaborate with cross-functional teams to design innovative data systems
- Leverage AI techniques to automate parsing of complex, dirty datasets
What they're looking for
- Python development (5-7 years professional experience)
- Data cleaning and transformation techniques
- Linux bash/ksh scripting and regular expressions
- AWS services (EC2, RDS, SQS, S3, Lambda, API Gateway)
- Relational and NoSQL databases (Elasticsearch, SQL, Databricks)
- ETL pattern design and automation
- Computer science fundamentals (data structures, algorithms)
- Git version control and CI/CD pipeline development
Benefits
- 401(k) with employer contribution
- Health, vision, and dental insurance with HSA option
- Employer-paid life, short-term, and long-term disability insurance
- Generous PTO and 16 paid holidays annually
- Flexible and remote-friendly work arrangements
- Engaging workspace in South Austin
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
SpyCloud
- Website
- spycloud.com
Likely interview questions
- Walk us through your experience building ETL pipelines in Python and handling unstructured data at scale.
- Describe a time you optimized a data processing workflow—what challenges did you face and how did you solve them?