Senior Elasticsearch Engineer owning full lifecycle of massive-scale search and analytics platform at Chess.com: capacity planning, architecture, performance tuning, incident response, and Elasticsearch-to-OpenSearch migrations on bare-metal Kubernetes. Requires 7+ years operating Elasticsearch at scale with deep internals knowledge.
Salary not listed
Remote7+ YOEDevOps / SRE
About the role
What you'll do
Own incident response and reliability for Elasticsearch/OpenSearch clusters: shard allocation strategy for write-heavy data streams (millions of documents per minute), disk watermark management, retention policy tuning, rollover orchestration, performance optimization, I/O tuning on bare-metal nodes, write queue analysis, thread pool diagnostics, shard rebalancing under load, capacity planning, and growth forecasting.
Provide on-call ownership for Elasticsearch-related incidents including cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades; perform real-time cluster triage, cross-team coordination, post-mortem authoring, and systemic reliability improvements.
Manage snapshot and disaster recovery across clusters.
Lead Elasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation for ILM/ISM, security models, and plugin ecosystems; handle version upgrade planning, rolling restart orchestration with zero-downtime, and end-to-end new cluster provisioning.
Advise engineering teams on index design, mapping strategy, retention policies, and query optimization; manage Kibana and OpenSearch Dashboards access/configuration; define and maintain workload priority tiers.
Requirements
7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput).
Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management.
Production experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator-based deployments.
Proficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management).
Hands-on Linux systems administration with focus on storage and I/O performance.
Experience managing both Elasticsearch and OpenSearch in production, with informed opinions on their trade-offs.
Incident command experience: diagnose and mitigate cluster emergencies under pressure while communicating clearly.
Senior engineer building and operating foundational Dev Tools infrastructure at Airbnb, including cloud development environments, Kubernetes platforms for AI agents, source control, code review, and polyglot monorepo tooling. Requires 5+ years building high-scale distributed systems with a focus on developer productivity, reliability, and cost efficiency.
196k – 230k/yr
Remote5+ YOEDevOps / SRE
Network Production Engineering Lead
FluidstackSan Francisco, CA
Lead the network production engineering team responsible for availability, performance, and automation of fabrics supporting 100k+ accelerator clusters at massive scale. Own SLOs, build remediation automation, and set operating models between design and site teams.
242k – 284k/yr
On-site7+ YOEDevOps / SRE
AI Enablement Engineer
Sprinter HealthSan Francisco, CA
Build and enable company-wide AI adoption by creating agents, workflows, prompt libraries, evaluation frameworks, and training programs. Partner with engineering, clinical, operations and other teams to identify opportunities, deliver production AI tools, and ensure safe, measurable impact in a regulated healthcare environment.
180k – 260k/yr
Hybrid7+ YOEDevOps / SRE
Senior Performance Engineer
CrusoeSan Francisco, CA
Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.
170k – 205k/yr
On-site5+ YOEDevOps / SRE
Senior Software Engineer
Grafana LabsUnited States
Senior SRE embedded with Mimir and Loki squads to own production reliability and SLOs for Grafana Cloud's high-SLA database products (Mimir, Loki, Tempo, Pyroscope) running on AWS/GCP/Azure. Requires 6+ years engineering experience including 3+ in SRE/production engineering, strong Kubernetes and multi-tenant systems experience, and on-call participation.