Palantir
Product Reliability Engineer - Defense
About this role
Palantir seeks a Product Reliability Engineer to ensure the health and performance of mission-critical software services used by defense and government partners. You'll own end-to-end service reliability, respond to outages, eliminate technical debt, and drive long-term stability improvements across complex systems.
What you'll do
- Respond to service outages and customer-reported issues during on-call shifts
- Investigate root causes and implement permanent solutions to reliability problems
- Introduce observability and monitoring into complex systems
- Execute infrastructure migrations and codebase modernization efforts
- Collaborate with product teams to improve system stability and resilience
- Address technical debt in essential codebases
What they're looking for
- Troubleshooting and debugging complex systems
- Infrastructure and systems engineering
- Observability and monitoring tool expertise
- On-call incident response
- Software architecture and design
- Technical documentation
- Collaboration and communication
- Ownership mentality and urgency
Benefits
- Mentorship from experienced engineers
- Structured onboarding program
- Work on mission-critical defense applications
- Opportunity to impact lifesaving initiatives
- Focus on both operational and strategic technical work
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Palantir
Palantir builds data platforms and software solutions that help government and enterprise customers tackle complex operational challenges, with a focus on responsible AI governance and privacy. The company is hiring software engineers and interns for forward-deployed customer roles, infrastructure and platform teams, and specialized privacy and civil liberties engineering positions.
- Website
- palantir.com
Likely interview questions
- Walk us through a time you debugged a production outage. How did you approach it, and what permanent solution did you implement rather than applying a quick fix?
- Describe your experience with observability tools and monitoring. How have you used metrics, logs, or traces to identify and prevent reliability issues?