Skip to main content

SpyCloud

Data Processing Engineer

Austin, Texas | Remote (Remote)$111k–$144kfull-timemidAdded yesterday

About this role

SpyCloud seeks a Security Data Analyst/Python Developer to parse, transform, and automate processing of large datasets for cybersecurity applications. This hybrid/remote role requires 5-7 years of Python experience, AWS proficiency, and expertise in data cleaning and ETL pipelines to support the company's mission of disrupting cybercrime.

What you'll do

  • Parse, clean, and transform structured and unstructured datasets to align with SpyCloud schema
  • Build Python-based automation for the data parsing platform and develop ETL scripts
  • Monitor data ingestion processes and troubleshoot issues
  • Develop CI/CD pipelines and establish standards for codebase management
  • Collaborate with cross-functional teams to design innovative data systems
  • Leverage AI techniques to automate parsing of complex, dirty datasets

What they're looking for

  • Python development (5-7 years professional experience)
  • Data cleaning and transformation techniques
  • Linux bash/ksh scripting and regular expressions
  • AWS services (EC2, RDS, SQS, S3, Lambda, API Gateway)
  • Relational and NoSQL databases (Elasticsearch, SQL, Databricks)
  • ETL pattern design and automation
  • Computer science fundamentals (data structures, algorithms)
  • Git version control and CI/CD pipeline development

Benefits

  • 401(k) with employer contribution
  • Health, vision, and dental insurance with HSA option
  • Employer-paid life, short-term, and long-term disability insurance
  • Generous PTO and 16 paid holidays annually
  • Flexible and remote-friendly work arrangements
  • Engaging workspace in South Austin
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

SpyCloud

View all jobs at SpyCloud

Likely interview questions

  • Walk us through your experience building ETL pipelines in Python and handling unstructured data at scale.
  • Describe a time you optimized a data processing workflow—what challenges did you face and how did you solve them?