Forward Deployed Reliability Engineer
Forward Deployed Reliability Engineer ensures stability of Palantir's mission-critical workflows by handling on-call incidents, automating solutions, and driving product improvements. Requires proficiency in Python, Java, SQL, and a technical background.
About the job
Core Responsibilities
- Develop a deep understanding of Palantir's products and operational processes
- Go on-call, responding quickly and effectively to mission-critical incidents
- Diagnose, resolve, and proactively prevent issues encountered in the field
- Collaborate with internal stakeholders to increase the scalability and reliability of Foundry workflows for our customers
- Identify recurring pain points and inefficiencies, and take initiative to automate or streamline workflows
- Advocate for and implement product enhancements based on insights gleamed from the field
- Create clear, actionable documentation and share best practices to elevate team and company-wide reliability
Note: While active work is not required on weekends or outside business hours, you must be available to respond to critical outages during assigned on-call weeks.
What We Value
- Ability to work independently and collaboratively to solve ambiguous technical and operational challenges
- Excellent written and verbal communication skills, capable of interacting effectively with both technical and non-technical stakeholders
- Proficiency in Python, Java, and SQL
- Familiarity with parallel data processing and Spark job optimization
- Strong organizational skills and attention to detail, with the ability to prioritize effectively
- Resourcefulness and creativity in fast-paced dynamic environments
- Experience with root cause analysis and documenting solutions for broader impact
- Enthusiasm for hands-on problem solving, continuous improvement, and knowledge sharing
What We Require
- Background in Computer Science, Engineering, Information Systems, or other technical field.
Skills
Python, Java, SQL, Spark, Foundry, Root Cause Analysis
Similar jobs
DevOps / SRE jobsDesigns and operates foundational developer-infrastructure services for CI, builds, deployments, and testing. The role requires senior-level systems engineering, end-to-end service ownership, and cross-functional technical leadership.
Build and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.