Skip to content
Chess.comChess.comUnited States

Senior Elasticsearch Engineer

Senior Elasticsearch Engineer owning full lifecycle of massive-scale search and analytics platform at Chess.com: capacity planning, architecture, performance tuning, incident response, and Elasticsearch-to-OpenSearch migrations on bare-metal Kubernetes. Requires 7+ years operating Elasticsearch at scale with deep internals knowledge.

Salary not listed
Remote7+ YOEDevOps / SRE

About the role

What you'll do

  • Own incident response and reliability for Elasticsearch/OpenSearch clusters: shard allocation strategy for write-heavy data streams (millions of documents per minute), disk watermark management, retention policy tuning, rollover orchestration, performance optimization, I/O tuning on bare-metal nodes, write queue analysis, thread pool diagnostics, shard rebalancing under load, capacity planning, and growth forecasting.
  • Provide on-call ownership for Elasticsearch-related incidents including cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades; perform real-time cluster triage, cross-team coordination, post-mortem authoring, and systemic reliability improvements.
  • Manage snapshot and disaster recovery across clusters.
  • Lead Elasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation for ILM/ISM, security models, and plugin ecosystems; handle version upgrade planning, rolling restart orchestration with zero-downtime, and end-to-end new cluster provisioning.
  • Advise engineering teams on index design, mapping strategy, retention policies, and query optimization; manage Kibana and OpenSearch Dashboards access/configuration; define and maintain workload priority tiers.

Requirements

  • 7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput).
  • Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management.
  • Production experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator-based deployments.
  • Proficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management).
  • Hands-on Linux systems administration with focus on storage and I/O performance.
  • Experience managing both Elasticsearch and OpenSearch in production, with informed opinions on their trade-offs.
  • Incident command experience: diagnose and mitigate cluster emergencies under pressure while communicating clearly.
  • Git-based infrastructure management (GitOps): Helm charts, ArgoCD/Flux, infrastructure-as-code.
  • Fluency with Elastic stack APIs: cluster administration, index templates, data streams, ILM policies, snapshot/restore.

Preferred Skills

  • OpenSearch ISM policies and security plugin (fine-grained access control).
  • GCS or S3 snapshot repository configuration and cross-cluster replication.
  • Grafana + Prometheus monitoring for Elasticsearch metrics.
  • Kibana Discover, Dev Tools, and data view management at scale.
  • Java internals relevant to Elasticsearch JVM tuning (heap sizing, GC tuning, circuit breakers).
  • Vault integration for secrets management in Kubernetes-deployed search clusters.
  • Fluentd/Fluent Bit log pipeline configuration feeding OpenSearch.
  • Hardware selection experience for search-optimized server configurations.
  • Python or scripting for operational analysis and automation.

Skills

ElasticsearchopensearchKuberneteseckGitOpsHelmArgo CDilmismkibanaPrometheusGrafanaJavaPythonLinux

Similar roles

DevOps / SRE jobs
Airbnb

Senior Software Engineer, Dev Tools

AirbnbUnited States

Senior engineer building and operating foundational Dev Tools infrastructure at Airbnb, including cloud development environments, Kubernetes platforms for AI agents, source control, code review, and polyglot monorepo tooling. Requires 5+ years building high-scale distributed systems with a focus on developer productivity, reliability, and cost efficiency.

196k – 230k/yr
Remote5+ YOEDevOps / SRE
Fluidstack

Network Production Engineering Lead

FluidstackSan Francisco, CA

Lead the network production engineering team responsible for availability, performance, and automation of fabrics supporting 100k+ accelerator clusters at massive scale. Own SLOs, build remediation automation, and set operating models between design and site teams.

242k – 284k/yr
On-site7+ YOEDevOps / SRE
Sprinter Health

AI Enablement Engineer

Sprinter HealthSan Francisco, CA

Build and enable company-wide AI adoption by creating agents, workflows, prompt libraries, evaluation frameworks, and training programs. Partner with engineering, clinical, operations and other teams to identify opportunities, deliver production AI tools, and ensure safe, measurable impact in a regulated healthcare environment.

180k – 260k/yr
Hybrid7+ YOEDevOps / SRE
Crusoe

Senior Performance Engineer

CrusoeSan Francisco, CA

Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.

170k – 205k/yr
On-site5+ YOEDevOps / SRE
Grafana Labs

Senior Software Engineer

Grafana LabsUnited States

Senior SRE embedded with Mimir and Loki squads to own production reliability and SLOs for Grafana Cloud's high-SLA database products (Mimir, Loki, Tempo, Pyroscope) running on AWS/GCP/Azure. Requires 6+ years engineering experience including 3+ in SRE/production engineering, strong Kubernetes and multi-tenant systems experience, and on-call participation.

154k – 185k/yr
Remote6+ YOEDevOps / SRE