Staff DevOps Engineer
Operates and improves large-scale, multi-region blockchain RPC infrastructure across Kubernetes, cloud, and bare-metal environments. The role requires production operations, infrastructure-as-code, GitOps, observability, incident response, and on-call experience.
About the job
Responsibilities
- Deploy, operate, and maintain blockchain RPC nodes across multiple chains and geographic regions.
- Manage Kubernetes clusters underlying the blockchain node platform.
- Perform rolling upgrades and hard fork migrations for blockchain clients across EVM and non-EVM chains.
- Participate in on-call rotations, triage live incidents via PagerDuty, and coordinate resolution for node outages, latency spikes, and SLO breaches.
- Develop and maintain AI agents and automation tooling for health checks, auto-healing, and hard fork notifications.
- Deploy and manage services using ArgoCD and GitOps workflows with Helm charts.
- Manage bare-metal and cloud infrastructure, including provisioning, benchmarking, and hardware replacement.
- Respond promptly to security advisories and coordinate upgrades with minimal downtime.
- Contribute to postmortems and asynchronous review processes; track action items and follow up on resolutions.
- Collaborate with product, customer success, and engineering teams on chain deprecations, capacity planning, and SLO reporting.
Requirements
- Experience designing and operating large-scale, multi-region, multi-cloud production systems.
- Experience with Kubernetes, including StatefulSets, storage management, Secrets, and service mesh technologies such as Istio.
- Experience with secrets management and access control in multi-cluster environments.
- Familiarity with automation frameworks for node health checks, upgrades, and remediation workflows.
- Experience with infrastructure as code, such as Terraform, Ansible, Pulumi, CloudFormation, Chef, or Puppet.
- Experience with GitOps tooling, including ArgoCD and Helm.
- Proficiency with cloud infrastructure and bare-metal management, including storage provisioning and snapshot management.
- Strong understanding of observability tooling, including Grafana, Prometheus, and Alertmanager, with experience building or tuning dashboards and alert rules.
- Comfort working in an on-call environment and triaging production incidents using PagerDuty and structured runbooks.
- Ability to write technical documentation and postmortems and contribute to asynchronous team communication.
- Experience with networking and configuring or managing VPC networks.
- Basic understanding of security best practices.
- Passion for blockchain technologies and Web3.
Nice-to-Haves
- Experience with service mesh deployments such as Istio.
- Good understanding of web applications and microservice architecture.
- Experience working with startups.
Compensation and Benefits
- Attractive salary package.
- Private medical insurance.
- Flexible time away.
- Internal off-site hackathons.
- Access to a company-rented hacker house during summer.
- Opportunities to travel across offices.
Skills
Kubernetes, Terraform, Ansible, Pulumi, CloudFormation, Argo CD, Helm, Istio, Grafana, Prometheus, Alertmanager, Pagerduty, GitOps, Vpc Networking, Infrastructure As Code
Similar jobs
DevOps / SRE jobsBuild and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.