Cursor
Software Engineer, Pretraining
San FranciscofulltimemidAdded today
About this role
Help build the data infrastructure powering frontier AI models. You'll work on large-scale web crawling, data quality systems, and pipeline platforms that transform raw internet-scale data into high-quality training datasets for coding AI.
What you'll do
- Design and operate high-throughput data pipelines that process frontier-scale data with full traceability and observability
- Build and train models for data classification, ranking, filtering, and quality assessment at extreme throughput
- Own web crawling and parsing systems including URL discovery, scheduling, host prioritization, and antibot handling
- Design scaling experiments to validate data mixtures and quality improvements against model loss and downstream evals
- Develop orchestration, tooling, and platform infrastructure to make data iteration fast, reliable, and reproducible
- Partner with training, acquisition, and research teams to turn data ideas into measurable capability improvements
What they're looking for
- Large-scale distributed systems and infrastructure design
- Data pipeline architecture and orchestration
- Web crawling and HTML parsing at scale
- Python or similar systems programming languages
- Machine learning model training and deployment
- Database and data warehouse systems
- Observability and monitoring systems
- Performance optimization and profiling
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Cursor
Cursor builds an AI-driven code editor used by millions of developers to transform how software is built. The company is hiring for infrastructure engineers, ML systems specialists, enterprise platform builders, security engineers, and customer success roles focused on driving adoption within large organizations.
- Website
- cursor.com
Likely interview questions
- Describe your experience building large-scale data pipelines—what were the biggest scaling challenges you faced?
- Tell us about a time you debugged a complex distributed system failure. How did you approach it?