Schonfeld
Site Reliability Engineer
About this role
Schonfeld seeks a Site Reliability Engineer to own the stability, scalability, and observability of their internal Agentic AI platform. You'll define reliability standards, manage incident response, support users directly, and continuously improve development practices in a fast-paced environment.
What you'll do
- Define and maintain Service Level Objectives (SLOs), error budgets, and incident response runbooks for the AI platform
- Own observability, incident response, reliability, and scalability across agents, gateways, LLM proxies, and RAG pipelines
- Investigate root causes, troubleshoot issues, and provide user support through a dedicated help channel
- Automate operational tasks and create monitoring workflows to ensure high availability and cost efficiency
- Review code quality, identify lifecycle gaps, and improve platform development standards and practices
- Monitor upstream dependencies and communicate status updates to stakeholders
What they're looking for
- Site Reliability Engineering or cloud automation (5+ years)
- Python application development
- AWS services (S3, OpenSearch, DynamoDB)
- Kubernetes and CI/CD pipelines (Github Actions)
- Relational and NoSQL databases (Postgres, MySQL, DynamoDB, Elasticsearch)
- REST API design and asynchronous architectures
- Problem-solving and technical communication
- Agentic AI, RAG pipelines, or observability tools (DataDog) — preferred
Benefits
- Competitive base salary ($175,000–$225,000)
- Performance bonus eligibility
- Comprehensive benefits package
- Learning and educational opportunities
- Collaborative, inclusive culture with internal networks and community initiatives
- Opportunities to make impactful contributions to a global hedge fund
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Schonfeld
Schonfeld builds internal platforms and technology systems that power investment operations, quantitative trading, and financial management for investment professionals. The company is hiring software engineers, platform engineers, and UI specialists to develop backend services, frontend solutions, data infrastructure, and trading platform support systems.
View all jobs at SchonfeldLikely interview questions
- Describe your experience defining and managing SLOs and error budgets—how did you balance reliability with development velocity?
- Tell us about a complex incident you investigated and resolved; walk us through your root cause analysis process.