Skip to main content

Windborne Systems

Machine Learning Infrastructure Engineer

  • Confirmed live in the last 24 hours
  • $140k–$240k
  • Mid level
  • Full-time
  • On-site · Palo Alto, California, United States
  • Added 2 months ago

About this role

WindBorne Systems is seeking a Machine Learning Infrastructure Engineer to build and maintain reliable infrastructure for their cutting-edge weather forecasting AI models. This role will focus on streamlining the research-to-operations pipeline, ensuring scalability and stability of their systems, and managing data pipelines from diverse sources. You'll play a critical role in enabling the research team to focus on model development and delivering timely, accurate weather intelligence.

What you'll do

  • Build and maintain research-to-operations pipelines with strict latency requirements.
  • Evaluate cost/performance tradeoffs across cloud and on-prem compute resources.
  • Develop data pipelines for training and real-time data, handling upstream delays and quality checks.
  • Improve the reliability of distributed training runs with monitoring and auto-recovery.
  • Manage and scale inference infrastructure to meet customer demands.
  • Troubleshoot and resolve production issues across nodes.

What they're looking for

  • PyTorch
  • Docker
  • Data Pipelines
  • Production ML Systems
  • Large Datasets
  • GPU Clusters
  • Job Schedulers
  • Cloud Computing

Benefits

  • 401(k)
  • Health, Dental, and Vision Insurance
  • Unlimited PTO
  • Stock Option Plan
  • Office Food & Beverages
  • Hybrid/In-Person Work
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Windborne Systems

View all jobs at Windborne Systems

Likely interview questions

  • Describe your experience building and maintaining production ML systems. What are some common challenges you’ve faced and how did you address them?
  • How do you approach building data pipelines that can handle unreliable or incomplete upstream data?