Skip to main content

Paradigm

Infrastructure Engineer, Applied AI

San Francisco, CA$250k–$400kfulltimemidAdded today

About this role

Paradigm seeks an Infrastructure Engineer to build and maintain the core systems powering Centaur, a self-hosted runtime for multiplayer AI agents. You'll own critical infrastructure components including control planes, sandboxed execution environments, security controls, and observability systems that directly enable the firm's AI research and investment operations.

What you'll do

  • Design and maintain control plane services that coordinate agent runs with durable state management and failover recovery
  • Build and harden Kubernetes-based sandboxed runtimes with isolation and network controls for safe untrusted code execution
  • Implement secure credential management enabling agents to call third-party APIs without exposing raw keys
  • Establish monitoring, alerting, incident response, and failure-recovery mechanisms for production reliability
  • Ensure Centaur remains self-hosted and deployable by external organizations with clean configuration and upgrade paths
  • Collaborate with research and investing teams to identify infrastructure gaps and experimental opportunities

What they're looking for

  • Kubernetes and container orchestration
  • Distributed systems and state management
  • Security and sandboxing techniques
  • Cloud infrastructure and DevOps
  • Monitoring and observability systems
  • Go, Rust, or Python backend development
  • API design and integration
  • Incident response and reliability engineering
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Paradigm

View all jobs at Paradigm

Likely interview questions

  • Describe your experience designing control planes or coordination systems for distributed workloads. What were the key durability challenges you faced?
  • How have you approached sandboxing and isolating untrusted code in production systems? What trade-offs did you consider?