Skip to main content

mthree Recruiting Portal

Reliability & Production Engineer

New York, NY$85k–$110kmidAdded today

About this role

Join a fintech team supporting mission-critical equities trading platforms and alternative trading systems. As a Reliability & Production Engineer in New York, you'll ensure platform stability, resolve production incidents, and drive continuous improvement across high-availability trading environments.

What you'll do

  • Monitor and maintain mission-critical trading applications and alternative trading systems
  • Resolve production incidents and troubleshoot complex distributed system issues
  • Ensure platform stability and high availability across the trading environment
  • Collaborate with traders, developers, and quantitative teams on operational challenges
  • Implement process improvements and system enhancements for production operations
  • Provide incident response and post-incident analysis to prevent future outages

What they're looking for

  • Production support and incident management
  • Distributed systems and high-availability architecture
  • Troubleshooting and root cause analysis
  • Electronic trading platforms knowledge
  • System monitoring and observability tools
  • Scripting or programming languages
  • Cross-functional communication and collaboration
  • Process improvement and automation
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

mthree Recruiting Portal

mthree Recruiting Portal connects recent graduates with entry-level technology roles at leading financial services and enterprise organizations through its Alumni program. The company specializes in recruiting, training, and placing new talent in Production Support and Site Reliability Engineer positions, with a focus on investment banking and other industries.

Website
mthree.com
View all jobs at mthree Recruiting Portal

Likely interview questions

  • Describe your experience supporting mission-critical production systems. How did you handle a major outage?
  • What tools and methodologies do you use for monitoring and alerting in high-availability environments?