Skip to main content

Cursor

Software Engineer, Pretraining

San FranciscofulltimemidAdded today

About this role

Help build the data infrastructure powering frontier AI models. You'll work on large-scale web crawling, data quality systems, and pipeline platforms that transform raw internet-scale data into high-quality training datasets for coding AI.

What you'll do

  • Design and operate high-throughput data pipelines that process frontier-scale data with full traceability and observability
  • Build and train models for data classification, ranking, filtering, and quality assessment at extreme throughput
  • Own web crawling and parsing systems including URL discovery, scheduling, host prioritization, and antibot handling
  • Design scaling experiments to validate data mixtures and quality improvements against model loss and downstream evals
  • Develop orchestration, tooling, and platform infrastructure to make data iteration fast, reliable, and reproducible
  • Partner with training, acquisition, and research teams to turn data ideas into measurable capability improvements

What they're looking for

  • Large-scale distributed systems and infrastructure design
  • Data pipeline architecture and orchestration
  • Web crawling and HTML parsing at scale
  • Python or similar systems programming languages
  • Machine learning model training and deployment
  • Database and data warehouse systems
  • Observability and monitoring systems
  • Performance optimization and profiling
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Cursor

Cursor builds an AI-driven code editor used by millions of developers to transform how software is built. The company is hiring for infrastructure engineers, ML systems specialists, enterprise platform builders, security engineers, and customer success roles focused on driving adoption within large organizations.

Website
cursor.com
View all jobs at Cursor

Likely interview questions

  • Describe your experience building large-scale data pipelines—what were the biggest scaling challenges you faced?
  • Tell us about a time you debugged a complex distributed system failure. How did you approach it?