Skip to main content

Together AI

Technical Support Engineer (GPU Clusters) - US Weekends

Remote (Remote)full-timemidAdded today

About this role

Support customers building AI solutions on Together AI's GPU clusters by resolving complex technical issues, monitoring infrastructure health, and serving as the primary technical expert before escalation to engineering. This full-time remote role works weekends and two weekdays after an initial weekday ramp-up period.

What you'll do

  • Resolve complex technical challenges for customers using Kubernetes GPU clusters and provide swift solutions
  • Monitor GPU cluster health and proactively alert customers to hardware issues with remediation steps
  • Operate production infrastructure including fleet rebalancing, Slurm maintenance, node repair, and workload management
  • Investigate and resolve storage and networking issues like Weka filesystem degradation and InfiniBand failures
  • Collaborate with Engineering and Product teams to address customer concerns and drive roadmap improvements
  • Create and maintain documentation, troubleshooting guides, and FAQs for internal and external knowledge sharing

What they're looking for

  • Kubernetes and container orchestration
  • GPU cluster management and troubleshooting
  • Linux/Unix system administration
  • Networking protocols (InfiniBand, Ethernet)
  • Storage systems (Weka or similar)
  • Slurm workload management
  • Customer support and communication
  • Infrastructure monitoring and observability
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Together AI

Together AI builds GPU compute infrastructure and open-source model customization platforms for AI developers and enterprises. The company is hiring for infrastructure operations, ML systems engineering, go-to-market technology, customer success, and GPU research roles.

View all jobs at Together AI

Likely interview questions

  • Walk us through your experience troubleshooting GPU cluster issues in a Kubernetes environment.
  • Describe a time you resolved a critical infrastructure issue under time pressure—what was your approach?