# Senior Elasticsearch Engineer

**Company:** [Chess.com](https://hotfix.jobs/companies/chess-com)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 7+ years
**Skills:** Elasticsearch, opensearch, Kubernetes, eck, GitOps, Helm, Argo CD, ilm, ism, kibana, Prometheus, Grafana, Java, Python, Linux
**Posted:** 2026-07-20

> Senior Elasticsearch Engineer owning full lifecycle of massive-scale search and analytics platform at Chess.com: capacity planning, architecture, performance tuning, incident response, and Elasticsearch-to-OpenSearch migrations on bare-metal Kubernetes. Requires 7+ years operating Elasticsearch at scale with deep internals knowledge.

## Job Description

## What you'll do
- Own incident response and reliability for Elasticsearch/OpenSearch clusters: shard allocation strategy for write-heavy data streams (millions of documents per minute), disk watermark management, retention policy tuning, rollover orchestration, performance optimization, I/O tuning on bare-metal nodes, write queue analysis, thread pool diagnostics, shard rebalancing under load, capacity planning, and growth forecasting.
- Provide on-call ownership for Elasticsearch-related incidents including cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades; perform real-time cluster triage, cross-team coordination, post-mortem authoring, and systemic reliability improvements.
- Manage snapshot and disaster recovery across clusters.
- Lead Elasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation for ILM/ISM, security models, and plugin ecosystems; handle version upgrade planning, rolling restart orchestration with zero-downtime, and end-to-end new cluster provisioning.
- Advise engineering teams on index design, mapping strategy, retention policies, and query optimization; manage Kibana and OpenSearch Dashboards access/configuration; define and maintain workload priority tiers.

## Requirements
- 7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput).
- Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management.
- Production experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator-based deployments.
- Proficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management).
- Hands-on Linux systems administration with focus on storage and I/O performance.
- Experience managing both Elasticsearch and OpenSearch in production, with informed opinions on their trade-offs.
- Incident command experience: diagnose and mitigate cluster emergencies under pressure while communicating clearly.
- Git-based infrastructure management (GitOps): Helm charts, ArgoCD/Flux, infrastructure-as-code.
- Fluency with Elastic stack APIs: cluster administration, index templates, data streams, ILM policies, snapshot/restore.

## Preferred Skills
- OpenSearch ISM policies and security plugin (fine-grained access control).
- GCS or S3 snapshot repository configuration and cross-cluster replication.
- Grafana + Prometheus monitoring for Elasticsearch metrics.
- Kibana Discover, Dev Tools, and data view management at scale.
- Java internals relevant to Elasticsearch JVM tuning (heap sizing, GC tuning, circuit breakers).
- Vault integration for secrets management in Kubernetes-deployed search clusters.
- Fluentd/Fluent Bit log pipeline configuration feeding OpenSearch.
- Hardware selection experience for search-optimized server configurations.
- Python or scripting for operational analysis and automation.

## Similar roles

- [Senior Software Engineer, Dev Tools](https://hotfix.jobs/jobs/dc931946-25e4-4987-a47b-f88c09268fd8) - Airbnb - Remote - $196k – $230k/yr
- [Network Production Engineering Lead](https://hotfix.jobs/jobs/51065552-0f77-47bc-a8f2-edf2c84caecc) - Fluidstack - San Francisco, CA - $242k – $284k/yr
- [AI Enablement Engineer](https://hotfix.jobs/jobs/307d7d5b-7858-4c37-bce0-624c01b780ac) - Sprinter Health - San Francisco, CA - $180k – $260k/yr
- [Senior Performance Engineer](https://hotfix.jobs/jobs/6629d9ee-877d-4387-8d8f-e75835fd70c0) - Crusoe - San Francisco, CA - $170k – $205k/yr
- [Senior Software Engineer](https://hotfix.jobs/jobs/fe40327f-1648-4f0e-ba97-f600e8bae0ef) - Grafana Labs - Remote - $154k – $185k/yr

**Apply:** https://hotfix.jobs/jobs/a4bac45a-192d-4739-a6fa-6bda2f1296dc
**Canonical:** https://hotfix.jobs/jobs/a4bac45a-192d-4739-a6fa-6bda2f1296dc