Skip to main content

Meridial

SWE-Bench AI Task Auditor - Freelance AI Trainer Project

World Wide - Remote (Remote)$145.6k–$208kcontractmidAdded today

About this role

Audit software engineering tasks designed to train and evaluate AI systems, ensuring they are technically accurate, realistic, and properly tested. Provide detailed feedback on codebase issues and evaluation criteria to maintain high-quality AI training workflows.

What you'll do

  • Evaluate software engineering tasks for technical accuracy, feasibility, and reproducibility
  • Assess test suites and evaluation criteria for reliability and completeness
  • Identify and troubleshoot codebase integration issues and logic errors
  • Provide clear, actionable feedback on identified technical problems
  • Verify tasks meet SWE-Bench standards and real-world engineering practices
  • Navigate complex codebases and validate solution approaches

What they're looking for

  • Software engineering expertise in your specialty domain
  • Complex codebase analysis and navigation
  • Test design and evaluation methodology
  • Problem-solving and debugging
  • Technical writing and documentation
  • Code review and quality assessment
  • Real-world application development experience
  • Understanding of AI training task requirements

Benefits

  • Flexible remote work from anywhere
  • Competitive hourly rate ($70-$100/hour)
  • Work on cutting-edge AI training projects
  • Leverage deep technical expertise
  • No company-mandated benefits (contractor arrangement)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Meridial

Meridial trains and improves advanced AI models through expert evaluation and feedback across infrastructure, software engineering, machine learning, and language domains. The company hires experienced freelance specialists—including software engineers, ML experts, and language specialists—to test AI reasoning, identify failure modes, and provide detailed training data to enhance model capabilities.

View all jobs at Meridial

Likely interview questions

  • Tell us about your experience navigating and analyzing complex codebases in production systems.
  • How do you approach identifying and evaluating whether a software task is technically sound and reproducible?