Incident Management Engineer
Leads response to critical product outages by triaging, troubleshooting, and coordinating resolutions across teams. Requires technical background in CS/Engineering, strong problem-solving, and comfort with 24/7 on-call in fast-paced environments.
About the job
Core Responsibilities
- Develop a deep understanding of Palantir’s product and delivery ecosystem.
- Collaborate with customer-facing, product, and infrastructure teams on the development and deployment of scalable, reliable software for our customers.
- Diagnose, resolve, and prevent issues encountered in the field.
- Reduce the operational overhead of responding to critical incidents at Palantir through investments in tooling, process, and automation.
- Take part in a 24/7 on-call rotation responsible for coordinating Palantir’s response to mission-critical incidents, ensuring efficient resolution with minimal customer impact.
What We Value
- Excellent problem solving skills.
- Comfort working in a fast paced environment.
- Ability to work both independently and make decisions under minimal direction, as well as collaborate as part of a team.
- Experience with scripting, automation, or data analysis a plus.
What We Require
- Background in Computer Science, Engineering, Information Systems, Incident Management, or other technical field.
- Willingness and interest to travel to other Palantir locations on occasion.
Skills
Incident Management, Scripting, Automation, Data Analysis, Troubleshooting, On-Call Rotation, Problem Solving
Similar jobs
DevOps / SRE jobsDesigns and operates foundational developer-infrastructure services for CI, builds, deployments, and testing. The role requires senior-level systems engineering, end-to-end service ownership, and cross-functional technical leadership.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.