Skip to main content

Latent Defense

Site Reliability Engineer

San Francisco$200k–$275kfulltimemidAdded 1 month ago

About this role

Own the production infrastructure for a clinical AI platform serving major health systems, ensuring 99.9%+ uptime while enabling rapid product development. You'll architect and maintain mission-critical systems using Kubernetes, Terraform, and modern DevOps practices in a high-intensity, in-office environment.

What you'll do

  • Design, implement, and maintain production environment for clinical AI infrastructure
  • Manage containerized infrastructure using Kubernetes and Helm at scale
  • Optimize CI/CD deployment pipelines for TypeScript and Python/ML applications
  • Define and implement Infrastructure as Code using Terraform
  • Support developer experience by streamlining workflows and tooling
  • Own operational excellence and establish standards for system reliability

What they're looking for

  • Kubernetes and Helm
  • Terraform and Infrastructure as Code
  • CI/CD pipeline optimization
  • Distributed systems architecture
  • Command line and automation
  • PostgreSQL, Redis, and Kafka
  • Deployment at scale (500+ machines)
  • Python and TypeScript
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Latent Defense

Latent Defense builds modern clinical AI systems designed to replace outdated EHRs and improve healthcare workflows for providers and patients. The company is hiring frontend engineers, DevOps infrastructure specialists, backend developers, and machine learning engineers to develop intuitive interfaces, ensure production reliability, power clinical data pipelines, and deploy personalized healthcare ML systems.

View all jobs at Latent Defense

Likely interview questions

  • Walk us through your experience managing large-scale infrastructure—how have you handled deployments at the 500+ machine scale, and what challenges did you face?
  • Describe your approach to designing and maintaining a Kubernetes cluster for a mission-critical system. How do you ensure 99.9%+ uptime?