Skip to content
PalantirPalantir

Production Engineer - Database Operations

This engineer automates deployment and operation of production databases, manages migrations, improves reliability, and participates in incident response and on-call rotations. The role requires programming experience, cloud and orchestration expertise, database knowledge, and a strong Linux foundation.

About the job

Responsibilities

  • Build software that automates the routine work of deploying and running production databases.
  • Participate in regular on-call rotations.
  • Troubleshoot, diagnose, and remediate stability and reliability issues in production database systems.
  • Participate in post-incident reviews and own follow-up actions.
  • Identify patterns across incidents, support tickets, and alerts and translate them into proposals for systemic fleet improvements.
  • Manage and execute large-scale migrations of the database fleet.
  • Partner with customer-facing teams during incidents and when setting up new database installations.
  • Work with database engineering and infrastructure teams to build resilient, highly available database systems.
  • Invest in documentation, metrics, monitors, and troubleshooting tools.
  • Maintain engineering and operational standards through code reviews, design reviews, and feedback on operating procedures.
  • Build and share expertise in production systems and technologies such as Kubernetes, cloud environments, Cassandra, and Elasticsearch.

Requirements

  • Engineering background in Computer Science, Mathematics, Software Engineering, Physics, or a similar field.
  • Experience writing code in languages such as Java, Python, and Go, through professional work or personal projects.
  • Experience with cloud and orchestration technologies such as Kubernetes, Helm, AWS, Google Cloud, Azure, and OpenShift.
  • Experience with database technologies such as Cassandra, Elasticsearch, and Kafka.
  • Solid foundation in Linux and distributed web services.
  • Familiarity with observability tools such as Grafana and Prometheus.
  • Strong written and verbal communication skills, with the ability to iterate quickly with teammates and incorporate feedback.

Skills

Java, Python, Go, Kubernetes, Helm, AWS, GCP, Azure, Openshift, Cassandra, Elasticsearch, Kafka, Linux, Grafana, Prometheus

Clickhouse

Clickhouse

Singapore
Release Engineer - Data Plane Internal Tooling and Productivity
No salary listedRemote5+ YOEDevOps / SRE

Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.

PostHog

PostHog

Remote

ClickHouse Operations Engineer
No salary listedRemoteDevOps / SRE

Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.