Senior DevOps Engineer
Leads the design, automation, and reliability of large-scale, multi-cloud infrastructure supporting search, NoSQL, and AI-driven workloads. Requires 7+ years in infrastructure, DevOps, or SRE, plus deep Kubernetes, Terraform, Linux, and distributed data-systems expertise.
About the job
Responsibilities
- Architect, scale, and maintain search and NoSQL technologies, including Solr, HBase, and Redis.
- Manage multi-cloud infrastructure across AWS and Google Cloud using Terraform.
- Deploy and operate stateful data services on Kubernetes, including stateful sets, persistent storage, and resource isolation.
- Lead infrastructure integration for vector search databases and high-performance computing supporting AI-driven architectures.
- Implement monitoring and alerting with Datadog and Prometheus.
- Automate operational tasks and maintain tooling in Python, Go, or Bash.
- Manage cloud networking, including VPCs, load balancing, service meshes, and caching layers.
- Debug complex distributed-systems incidents, conduct root-cause analysis, and participate in an on-call rotation.
Requirements
- 7+ years of experience in infrastructure, DevOps, or SRE, including substantial experience managing high-scale distributed data systems.
- Direct experience with Solr, HBase, and Redis, or comparable technologies such as Elasticsearch, OpenSearch, Cassandra, or Bigtable.
- Expert knowledge of Linux performance tuning for data-intensive applications, including I/O scheduling, memory management, and JVM tuning.
- Experience managing large-scale infrastructure with Terraform or OpenTofu.
- Deep experience running production-grade stateful workloads on Kubernetes.
- Strong proficiency in Python or Go, with the ability to build custom tools and operators.
- Practical experience using LLM tools such as GitHub Copilot, ChatGPT, or Claude for engineering productivity.
- Ability to learn new technologies independently and communicate complex technical problems clearly in English.
- Strong technical, organizational, interpersonal, written, and verbal communication skills.
Nice-to-haves
- Experience transitioning between comparable search or NoSQL technologies.
- Experience integrating vector search databases and high-performance computing into AI architectures.
- Experience with multi-cloud environments and service meshes.
Skills
Solr, Hbase, Redis, Terraform, AWS, GCP, Kubernetes, Datadog, Prometheus, Python, Go, Bash, Linux, Jvm, Opentofu
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.
Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.
Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.