Skip to main content

Crusoe

Staff Product Manager, AI Infrastructure (Fleet Management)

  • Confirmed live in the last 24 hours
  • No salary listed
  • Senior and above
  • Full-time
  • On-site · San Francisco, CA - US
  • 6+ yrs exp
  • Added today

About this role

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role

The Staff Product Manager, Fleet Management owns the product strategy and execution for the systems that keep Crusoe's GPU and CPU fleet running reliably and efficiently at scale. You'll define how Crusoe provisions, monitors, maintains, and optimizes its server fleet supporting AI training and inference workloads, translating operational requirements into scalable product capabilities. This covers the full lifecycle of compute: initial provisioning, firmware and configuration management, burn-in, multi-node testing, health monitoring, maintenance orchestration, repair, and decommissioning.

The role sits at the intersection of infrastructure operations, engineering, and customer experience. Fleet-level systems determine how much of our capacity is sellable, how fast we can bring new clusters online, and how quickly a customer's workload recovers when hardware fails. You'll be defining the problem space as much as solving it, and you should be comfortable operating in ambiguity, with support from senior leadership as the initiative takes shape.

You'll collaborate closely with hardware partners, engineering, infrastructure operations, networking, supply chain, finance, and customer success to ensure cohesive lifecycle management across Crusoe Cloud. As a Staff PM, you'll lead significant cross-functional initiatives, influence technical and product strategy across multiple teams, and be accountable for outcomes that directly impact company goals.

What You'll Be Working On

  • Own the vision and product strategy for fleet management, including provisioning, burn-in, lifecycle management, health monitoring, maintenance orchestration, and performance optimization

  • Drive outcomes across the entire product lifecycle, from discovery through launch and iteration, for fleet management capabilities

  • Define product direction for systems managing large-scale GPU and CPU infrastructure, ensuring reliability, utilization, and operational efficiency

  • Lead cross-functional initiatives spanning engineering, infrastructure operations, networking, and customer success to deliver integrated fleet management solutions

  • Build consensus with tech leads and engineering managers on architecture decisions, tooling investments, and operational process improvements

  • Identify opportunities to improve fleet utilization, reduce operational overhead, and enhance observability through product innovation

  • Translate customer-facing reliability and performance impact, gathered through customer success, into product requirements

  • Mentor other product managers and engineers, sharing expertise in infrastructure operations and complex systems thinking

What You'll Bring to the Team

  • 6+ years of product management experience delivering infrastructure or platform products, with demonstrated ownership of complex, cross-functional initiatives

  • Deep expertise in fleet management concepts including provisioning, lifecycle management, monitoring, and operations at scale

  • Strong understanding of distributed systems, server hardware, and infrastructure automation

  • Proven ability to solve ambiguous, novel problems requiring research, invention, and creative analysis

  • Experience driving product vision for complex systems with long-term strategic impact, balancing technical tradeoffs and operational requirements

  • Track record of building consensus across engineering, operations, and infrastructure teams on technical direction and investment priorities

  • Strong analytical skills interpreting operational metrics, performance data, and utilization patterns to guide product decisions

  • Excellent communication skills articulating technical concepts and connecting them to business outcomes

  • Experience mentoring colleagues in product management or technical roles

Bonus Points

  • Experience with large-scale GPU infrastructure or AI/ML platform operations

  • Familiarity with infrastructure automation tools, orchestration systems (Kubernetes, Slurm), or configuration management platforms

  • Experience with observability platforms, monitoring systems, or incident management workflows

  • Knowledge of hardware lifecycle management, predictive maintenance, or capacity planning systems

Benefits

  • Competitive compensation

  • Restricted Stock Units

  • Paid time off & paid holidays

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

Compensation Range

Compensation will be paid in the range of up to $215,000-$260,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Written by Crusoe. Original job post

Skills mentioned

  • Kubernetes
  • Machine Learning
  • Node.js
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Crusoe

Crusoe builds AI infrastructure and data center systems, including cloud platforms, modular facilities, and manufacturing operations. The company is hiring engineers across software, mechanical engineering, facilities management, instrumentation, and CNC programming to design, optimize, and operate its mission-critical infrastructure.

Industry
Technology & Software
View all jobs at Crusoe

Preparing likely interview questions for this role…