Palantir
Software Engineer - Hosted Model Infrastructure
About this role
Palantir seeks a Software Engineer to build infrastructure enabling ML models in production across diverse environments including air-gapped networks, edge devices, and enterprise systems. You'll own end-to-end services spanning inference engines, GPU scheduling, deployment pipelines, and observability to deliver AI capabilities to mission-critical customers.
What you'll do
- Design and maintain services for deploying ML models in constrained and air-gapped environments
- Develop inference engines, GPU scheduling systems, and deployment pipelines
- Build observability and monitoring solutions for model operations
- Integrate ML infrastructure with Palantir's broader platform
- Ensure continuous testing and reliable delivery of model updates
- Support customers across defense, government, and enterprise sectors
What they're looking for
- Software engineering and full-stack development
- Machine learning infrastructure and MLOps
- GPU computing and resource optimization
- Deployment pipeline and CI/CD systems
- Observability and monitoring tools
- Systems design and distributed systems
- Experience with air-gapped or edge environments
- Python or systems programming languages
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Palantir
Palantir builds data platforms and software solutions that help government and enterprise customers tackle complex operational challenges, with a focus on responsible AI governance and privacy. The company is hiring software engineers and interns for forward-deployed customer roles, infrastructure and platform teams, and specialized privacy and civil liberties engineering positions.
- Website
- palantir.com
Likely interview questions
- Walk us through your experience deploying ML models to production. What were the key challenges around reproducibility, versioning, or dependency management?
- How have you approached GPU resource scheduling or optimization in constrained environments? What trade-offs did you have to consider?