Arena Intelligence
Infrastructure Engineer, TL
About this role
Arena Intelligence seeks an Infrastructure Engineer to build core systems powering real-world AI model evaluation at scale. You'll design and operate low-latency APIs, gateways, and streaming infrastructure that route traffic across multiple frontier models while handling unpredictable load and ensuring reliability.
What you'll do
- Design and implement low-latency, high-reliability APIs for leaderboards, models, and evaluation arenas
- Handle streaming responses and partial failure recovery across heterogeneous LLM providers
- Build enterprise-grade features including rate limiting, authentication, usage metering, and compliance
- Instrument infrastructure with distributed tracing, latency analysis, and real-time observability dashboards
- Integrate infrastructure with core evaluation platform and collaborate with research team
- Contribute to backend systems for leaderboards and evaluation platforms across the stack
What they're looking for
- Backend engineering and distributed systems design
- Go and/or Rust programming
- High-throughput API and gateway/proxy development
- LLM provider APIs (OpenAI, Anthropic, Google) and streaming protocols
- AWS or GCP, Kubernetes, Terraform, Postgres, and Redis
- Product-oriented thinking and developer experience design
- API gateway architecture (Kong, Envoy, custom systems)
Opens the official application on the employer’s site. No login required.
Arena Intelligence
Arena Intelligence builds a platform for evaluating AI model performance through real-world testing and human preference data, powering insights for AI labs and enterprises. The company is hiring data engineers, security engineers, machine learning scientists, and customer-facing technical roles to scale its evaluation infrastructure, protect against misuse, and deliver custom solutions to customers.
View all jobs at Arena IntelligenceLikely interview questions
- Describe your experience building high-throughput API gateways or proxy systems—what were the hardest scaling challenges?
- How have you handled streaming responses from multiple upstream providers with different failure modes and latency profiles?