Skip to content
PalantirPalantirNew York, NY

Product Reliability Engineer - Defense

Product Reliability Engineers own end-to-end service health for Palantir's critical platforms, combining on-call incident response with forward-looking work on observability, resilience, code improvements, and infrastructure migrations. Requires strong backend coding (Java), troubleshooting skills, and ownership in dynamic environments; US security clearance eligibility needed.

Salary not listed
HybridDevOps / SRE

About the role

Core Responsibilities

  • Continuously invest in documentation, metrics, monitors and other troubleshooting tools.
  • Participate in on-call rotations during business hours and occasional weekends to remediate the most pressing issues across the Palantir fleet.
  • Diagnose, resolve, and prevent issues encountered in the field. Deliver end-to-end improvements to core products based on field issues.
  • Improve observability by refactoring codepaths and introducing telemetry.
  • Identify and implement data-driven opportunities for improved service resilience.
  • Develop strategic opinions on stability investments and inform the vision for long-term product stability.

What We Value

  • Comfortable with and curious about large scale production systems and technologies (e.g. load balancing, monitoring, distributed systems, and configuration management).
  • Confidence in troubleshooting complex issues independently using observability tools and stack traces.
  • Familiarity with monitoring tools such as Prometheus and health checks.
  • Experience coding with Java, Go and/or web technologies (e.g. HTML, CSS, JavaScript, Python/Ruby, Django/Flask/Ruby on Rails, etc.) is a plus.
  • Track record of identifying bugs in codebases and contributing fixes leading to long term service stability.
  • Demonstrated ability making data-driven decisions and engaging with stakeholders on strategy.

What We Require

  • Engineering background in Computer Science, Mathematics, Software Engineering, Physics or similar field.
  • Ability to work with a high degree of ownership and a strong sense of urgency in a dynamic environment.
  • Experience producing code in backend languages such as Java, as part of a past role or personal projects.
  • Familiarity with storage and data processing systems and cloud infrastructure.
  • Strong written and verbal communication and ability to iterate quickly with teammates and incorporate feedback.
  • Eligibility and willingness to obtain a US Security clearance.

Skills

JavaGoPrometheusDistributed SystemsObservabilityMonitoringload balancingconfiguration managementCloud InfrastructuretelemetryPythonJavaScript

Similar roles

DevOps / SRE jobs
Elicit

Infrastructure Engineer

ElicitOakland, CA

Own and evolve Elicit's cloud infrastructure platform (AWS/GCP, Kubernetes, Terraform) to support scalable single-tenant enterprise deployments. Build observability, compliance (SOC 2), cost optimization, and developer experience while contributing to backend systems where infra meets application logic. Requires 5+ years infrastructure/SRE experience, strong Terraform and K8s expertise, and enthusiasm for AI coding agents.

Salary not listed
On-site5+ YOEDevOps / SRE
OpenAI

Simulation Environments Engineer

OpenAISan Francisco, CA

Build and maintain CI/CD pipelines, orchestration, and automation for large-scale robotics simulation (SIL/HIL) to support model training, evaluation, and RL workloads at OpenAI. Requires strong infra, distributed systems, and Python/C++/Rust experience.

230k – 385k/yr
Hybrid5+ YOEDevOps / SRE
Airbnb

Operations Engineer, BizTech

AirbnbUnited States

Operations Engineer using AI, LLMs, and intelligent automation to triage tickets, accelerate incident response, build self-healing observability, and automate repetitive operational work in Airbnb's BizTech Global Operations team.

136k – 160k/yr
Remote3+ YOEDevOps / SRE
OpenAI

Systems Integration Engineer, Build Systems | Consumer Devices

OpenAISan Francisco, CA

Build and evolve Bazel, Yocto, and Buildkite-based CI systems for OpenAI consumer device software. Focus on hermetic builds, remote caching, test optimization, observability, and AI-powered failure analysis to accelerate reliable shipping. Requires 5+ years building developer infrastructure at scale.

293k – 325k/yr
Hybrid5+ YOEDevOps / SRE
Airbnb

Software Engineer, CI Platform Infrastructure

AirbnbUnited States

Build and optimize a next-generation CI platform infrastructure for workflow orchestration, scheduling, caching, and autoscaling to accelerate software development for engineers and AI coding agents at scale. Requires interest in distributed systems and knowledge of Kubernetes, EC2, Golang, and Docker.

162k – 190k/yr
RemoteDevOps / SRE