Skip to main content

Figure

Helix AI Engineer, Pretraining

San Jose, CAmidAdded 1 month ago

About this role

Figure AI is seeking a Helix AI Engineer to develop large-scale foundation models that power autonomous humanoid robots. You'll design pretraining strategies across multimodal data and build distributed training systems that enable robots to perceive, reason, and act in the real world.

What you'll do

  • Design and train large-scale multimodal foundation models using text, vision, and robot-collected data
  • Develop pretraining strategies to improve generalization and transfer to embodied AI tasks
  • Implement and explore transformer-based and emerging foundation model architectures
  • Build and optimize distributed training pipelines across multi-node GPU clusters
  • Collaborate with video, generative, agent, and robot learning teams to integrate models into the autonomy stack
  • Design evaluation frameworks and contribute to post-training approaches like fine-tuning and alignment

What they're looking for

  • Large-scale foundation model training experience
  • Deep learning architectures, especially transformers
  • Distributed training and optimization at scale
  • Python and PyTorch proficiency
  • Experimental rigor and model design iteration
  • Scalable software engineering practices
  • Multimodal pretraining (bonus)
  • Scaling laws and dataset curation (bonus)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Figure

Figure develops advanced humanoid robots powered by AI technology. The company is hiring engineers across mechanical design, firmware development, manufacturing, quality assurance, and security to build and refine its autonomous robotic systems.

View all jobs at Figure

Likely interview questions

  • Walk us through your experience training large-scale foundation models. What was the scale (parameters, data), and what specific pretraining objectives or strategies did you use?
  • How have you approached designing and optimizing distributed training pipelines across multi-node GPU clusters? What challenges did you encounter and how did you solve them?