GPTZero
Investigations Engineer - NYC
About this role
GPTZero is seeking an Investigations Engineer to build data pipelines and research systems that uncover AI-generated content and hallucinations at scale, turning findings into compelling front-page investigations published by major media outlets. You'll combine software engineering with investigative journalism to detect and report on problematic AI usage across the web and publishing.
What you'll do
- Deploy AI detection models to identify compelling examples of AI-generated content and hallucinations on the internet
- Build web crawlers, search query generators, and targeted data collection systems to gather documents at scale
- Develop data processing pipelines and automations to monitor social feeds and generate investigative leads
- Create ML ranking and classification systems to detect fabricated or AI-generated content in documents
- Synthesize data findings into written reports, visualizations, and interactive experiences for publication
- Improve model performance by analyzing results across new domains and refining detection approaches
What they're looking for
- Python (2+ years for resilient data pipelines and long-running jobs)
- SQL and database management (PostgreSQL experience preferred)
- Web scraping and large-scale data crawling techniques
- Data analysis and visualization tools
- AWS or cloud infrastructure basics
- Machine learning ranking and classification methods
- Web development and interactive UI creation
- OSINT and evidence verification techniques
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
GPTZero
GPTZero builds AI detection and agentic systems designed to promote information integrity online. The company is hiring LLM Engineers to develop and optimize language models, multi-agent workflows, and detection tools used by millions globally.
- Website
- gptzero.me
Likely interview questions
- Walk us through a past project where you built a data pipeline at scale—what challenges did you face and how did you ensure reliability?
- Describe your experience with web scraping and crawling. How have you handled scale, rate limiting, and data quality issues?