Skip to main content

Prophetic Technologies

AI Evals Engineer — Evaluation Datasets & Ground Truth

  • Confirmed live in the last 24 hours
  • No salary listed
  • Mid level
  • Full-time
  • US (Remote)
  • Added 4 weeks ago

About this role

Prophetic Software is seeking an AI Evaluations Engineer to develop and maintain high-quality evaluation datasets for their AI-powered real estate platform. This role focuses on defining 'correctness', sourcing reliable data (through human labeling, synthetic generation, or programmatic methods), and ensuring data trustworthiness over time. You'll work closely with product and engineering teams to close the feedback loop and contribute to continuous improvement of the platform.

What you'll do

  • Define evaluation criteria and rubrics for various modules.
  • Source data through production sampling, human labeling, or synthetic methods.
  • Version datasets and track lineage to ensure data integrity.
  • Calibrate automated grading systems against human-labeled gold sets.
  • Partner with engineers to integrate datasets into evaluation pipelines.

What they're looking for

  • ML/LLM Evaluation
  • Python
  • SQL
  • Prompt Engineering
  • Human Labeling
  • System Design
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Prophetic Technologies

View all jobs at Prophetic Technologies

Likely interview questions

  • Describe a time you built an evaluation dataset and what challenges you faced.
  • How do you ensure data integrity and prevent contamination in evaluation sets?