Skip to main content

Kled AI

Research Scientist, Benchmarks & Evals

  • Confirmed live in the last 24 hours
  • $150k–$250k
  • Mid level
  • Full-time
  • Remote · United States
  • 3+ yrs exp
  • Added today

About this role

ABOUT KLED

Kled is building the largest opt-in human data network in the world.

We are not a labeling firm. We are not a task marketplace.

We are a consumer application where people upload their real photos, videos, and documents and get paid continuously.

We then filter, standardize, and license that data to frontier AI labs and enterprises that need fresh, rights-aware training data.

Since launching our mobile app in 2026, we have:

• Reached #1 on the App Store (Finance) with 0 paid marketing
• Scaled to 200,000+ active data contributors
• Processed 1.5–3M uploads per day
• Raised $5M+ from investors behind SpaceX, Airbnb, Coinbase, xAI, OpenAI, Anthropic, Spotify, Lyft, Uber, and more

Our mission is to let anyone download the app and earn a real living wage from uploading their data.

ABOUT THE ROLE

Research Scientist - Benchmarks & Evals

AI labs buy data when a benchmark shows it works.

You’ll build the benchmarks that show what one of the largest human data networks in the world can do for frontier AI models.

You will:

- Design and ship public benchmarks for image, video and document models

- Build private test sets from real, opt-in data that no model has trained on

- Run data-value experiments: train models with and without Kled data, measure what changes

- Design human evaluation studies and the automatic metrics that track them

- Publish leaderboards and reports that researchers at frontier labs trust

- Find where models fall short and tell our team what data to collect next

- Work directly with the founders and lab buyers to turn results into data deals

WE’RE LOOKING FOR

- 3+ years ML research or research engineering experience

- A benchmark or evaluation you built that other people actually used

- Strong hands-on experience training and fine-tuning image or video models (PyTorch)

- Deep understanding of experimental design and statistics (controls, ablations, human studies)

- Experience running experiments end to end, from GPUs to final report

- Clear writing for both researchers and buyers

Bonus:

- Publications at NeurIPS, ICML, ICLR or CVPR (datasets and benchmarks especially)

- Experience on an evaluation or data team at an AI lab

- Experience selling or licensing data to AI labs

- Experience with data valuation or data attribution research

CURRENT STACK

Research

- Python / PyTorch

- Open-weight image and video models

- Cloud GPUs

Data

- Hundreds of millions of opt-in photos, videos and documents

- PostgreSQL (Supabase)

- S3 storage

COMPENSATION

- Base salary: $150,000 - $250,000

- $350,000 – $750,000 equity

We move fast and work hard.

If you're excited to build the benchmarks that decide what data trains frontier AI, let’s talk!

GROWTH OPPORTUNITY

You’ll join a team operating at the frontier of applied AI data infrastructure.

In this role, you’ll have the opportunity to:

• Own core systems that power one of the largest human data networks in the world
• Design infrastructure that directly influences what data trains next-generation AI models
• Build at real scale - millions of uploads per day, adversarial environments, global contributors
• Ship alongside a team that has built marketplaces, AI systems, and products used by millions

If you’re excited to move fast, build systems that matter, and help define how human data powers frontier AI, let’s talk.

Written by Kled AI. Original job post

Skills mentioned

  • Machine Learning
  • PostgreSQL
  • Python
  • PyTorch
  • S3
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Kled AI

View all jobs at Kled AI

Likely interview questions

  • Describe a benchmark or evaluation you've built that was used by others, and what impact did it have?
  • How would you approach designing a public benchmark for image, video, or document models, considering Kled's unique dataset?