Skip to content
Chess.comChess.com

Senior Elasticsearch Engineer

Senior Elasticsearch Engineer owning full lifecycle of massive-scale search and analytics platform at Chess.com: capacity planning, architecture, performance tuning, incident response, and Elasticsearch-to-OpenSearch migrations on bare-metal Kubernetes. Requires 7+ years operating Elasticsearch at scale with deep internals knowledge.

About the job

What you'll do

  • Own incident response and reliability for Elasticsearch/OpenSearch clusters: shard allocation strategy for write-heavy data streams (millions of documents per minute), disk watermark management, retention policy tuning, rollover orchestration, performance optimization, I/O tuning on bare-metal nodes, write queue analysis, thread pool diagnostics, shard rebalancing under load, capacity planning, and growth forecasting.
  • Provide on-call ownership for Elasticsearch-related incidents including cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades; perform real-time cluster triage, cross-team coordination, post-mortem authoring, and systemic reliability improvements.
  • Manage snapshot and disaster recovery across clusters.
  • Lead Elasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation for ILM/ISM, security models, and plugin ecosystems; handle version upgrade planning, rolling restart orchestration with zero-downtime, and end-to-end new cluster provisioning.
  • Advise engineering teams on index design, mapping strategy, retention policies, and query optimization; manage Kibana and OpenSearch Dashboards access/configuration; define and maintain workload priority tiers.

Requirements

  • 7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput).
  • Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management.
  • Production experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator-based deployments.
  • Proficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management).
  • Hands-on Linux systems administration with focus on storage and I/O performance.
  • Experience managing both Elasticsearch and OpenSearch in production, with informed opinions on their trade-offs.
  • Incident command experience: diagnose and mitigate cluster emergencies under pressure while communicating clearly.
  • Git-based infrastructure management (GitOps): Helm charts, ArgoCD/Flux, infrastructure-as-code.
  • Fluency with Elastic stack APIs: cluster administration, index templates, data streams, ILM policies, snapshot/restore.

Preferred Skills

  • OpenSearch ISM policies and security plugin (fine-grained access control).
  • GCS or S3 snapshot repository configuration and cross-cluster replication.
  • Grafana + Prometheus monitoring for Elasticsearch metrics.
  • Kibana Discover, Dev Tools, and data view management at scale.
  • Java internals relevant to Elasticsearch JVM tuning (heap sizing, GC tuning, circuit breakers).
  • Vault integration for secrets management in Kubernetes-deployed search clusters.
  • Fluentd/Fluent Bit log pipeline configuration feeding OpenSearch.
  • Hardware selection experience for search-optimized server configurations.
  • Python or scripting for operational analysis and automation.

Skills

Elasticsearch, Opensearch, Kubernetes, Eck, GitOps, Helm, Argo CD, Ilm, Ism, Kibana, Prometheus, Grafana, Java, Python, Linux

Shield AI

Shield AI

San Diego, CA
Senior Platform Engineer
$141k+/yrHybrid7+ YOEDevOps / SRE

Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.

Shield AI

Shield AI

San Mateo, CA
Senior Network Engineer
$140k+/yrOn-site6+ YOEDevOps / SRE

Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.

Astra

Astra

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.

Coinbase

Coinbase

United States

Senior Software Engineer, Core Infra Systems
$186k+/yrRemote5+ YOEDevOps / SRE

Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.

Shield AI

Shield AI

Seattle, WA
Senior Site Infrastructure Engineer
$110k+/yrOn-site5+ YOEDevOps / SRE

Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.