Skip to main content

Persona

Software Engineer, Resilience

San Francisco (Remote)$130k–$220kfulltimemidAdded 1 month ago

About this role

Persona is seeking a Software Engineer for their new Resilience Engineering function. This role focuses on enhancing performance and scalability by partnering with product teams and driving improvements across the organization's systems.

What you'll do

  • Collaborate with product teams on performance and observability projects
  • Analyze major incidents to identify architectural improvements
  • Design and promote reusable resilience tools and patterns
  • Advance the observability strategy and develop necessary tools
  • Mentor engineers to enhance performance and reliability skills
  • Establish an Operational Excellence review process

What they're looking for

  • 7+ years of software engineering experience
  • Performance, scalability, or observability knowledge
  • Mentorship abilities
  • Experience with observability tools
  • Empathy and trust-building skills
  • Capability to create robust libraries and tools
  • Initiative and comfort with ambiguity

Benefits

  • Medical, dental, and vision insurance
  • 3% 401(k) contribution
  • Unlimited PTO
  • Quarterly mental health days
  • Professional development stipend
  • Wellness benefits
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Persona

Persona builds an identity verification platform supported by data products, infrastructure, and networking systems that handle large-scale operations. The company is hiring Software Engineers for data products and platform teams, a networking specialist, a Security Engineer, and Solutions Engineers to support customer implementations and platform reliability.

View all jobs at Persona

Likely interview questions

  • Tell us about a time you diagnosed and resolved a critical performance or scalability issue in a production system. What was your approach, and how did you prevent similar issues in the future?
  • Describe your experience with observability — specifically metrics, logs, and traces. How have you used these to instrument and debug a complex distributed system?