Skip to main content

Palantir

Forward Deployed Reliability Engineer

New York, NYfull-timemidAdded 1 month ago

About this role

Palantir seeks a Forward Deployed Reliability Engineer to ensure stability of mission-critical workflows on their data platform. You'll respond to on-call incidents, resolve issues proactively, and drive product improvements by translating operational learnings into systemic enhancements and documentation.

What you'll do

  • Respond to on-call alerts and resolve mission-critical issues before customer impact
  • Diagnose and troubleshoot problems with Palantir software deployments and workflows
  • Automate manual tasks and implement technical workarounds to improve reliability
  • Document findings and create best practices guides for the broader team
  • Advocate for product and operational improvements based on field experience
  • Collaborate with customers and internal teams to refine processes and tooling

What they're looking for

  • On-call incident response and troubleshooting
  • Software reliability and operational excellence
  • Scripting and automation
  • Problem-solving and creative thinking
  • Technical documentation
  • Customer engagement and communication
  • System design and resilience
  • Data platform experience
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Palantir

Palantir builds data platforms and software solutions that help government and enterprise customers tackle complex operational challenges, with a focus on responsible AI governance and privacy. The company is hiring software engineers and interns for forward-deployed customer roles, infrastructure and platform teams, and specialized privacy and civil liberties engineering positions.

View all jobs at Palantir

Likely interview questions

  • Describe a time you responded to a critical production issue. Walk us through your debugging process and how you balanced immediate remediation with finding the root cause.
  • How do you approach documenting incident learnings and turning them into actionable improvements for your team or product?