TensorWave
Technical Support Engineer
About this role
TensorWave seeks a Technical Support Engineer to build their level 2 support team from the ground up, handling complex escalations from AI companies running GPU-intensive workloads at scale. You'll diagnose issues across Linux, Kubernetes, Slurm, and GPU infrastructure, drive tickets to resolution, and establish runbooks and diagnostic tools that scale the support operation.
What you'll do
- Own and resolve level 2 escalated tickets from the global operations center, or provide clean handoffs to engineering with full evidence
- Diagnose failures across Linux systems, GPU health, Kubernetes clusters, Slurm scheduling, and high-speed networking using logs and telemetry
- Investigate degraded or failed training and inference workloads, providing customers with root cause analysis and clear explanations
- Triage GPU and node hardware faults, identify RMA candidates, and coordinate remediation with data center operations
- Write and maintain runbooks to enable level 1 resolution of recurring issues; build diagnostic scripts to reduce investigation time
- Participate in on-call rotation, communicate directly with customer engineering teams, and feed patterns back to product and engineering teams
What they're looking for
- Linux systems administration and deep troubleshooting (kernel, networking, file systems, storage, process investigation)
- Kubernetes cluster diagnostics and workload scheduling understanding
- Slurm batch scheduler experience in shared compute environments
- Python or Bash scripting for automation and diagnostics
- GPU infrastructure and hardware fault diagnosis
- Log analysis, metrics interpretation, and telemetry investigation
- Incident and ticket tracking systems (JIRA, PagerDuty, or similar)
- Technical communication under pressure with senior engineering teams
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
TensorWave
TensorWave builds infrastructure and platforms for large-scale GPU clusters and AI workloads, spanning cloud, data center construction, and virtualization environments. The company is hiring Senior Software Engineers, DevOps Engineers, cloud infrastructure specialists, and operations engineers to develop automation tools, manage complex infrastructure systems, and support high-performance computing across multiple environments.
View all jobs at TensorWaveLikely interview questions
- Walk us through your most complex Linux troubleshooting investigation—what tools did you use and how did you isolate the root cause?
- Describe your experience with Kubernetes; have you diagnosed scheduling issues or workload placement problems in production?