Skip to content
AlchemyAlchemy

Staff DevOps Engineer

Operates and improves large-scale, multi-region blockchain RPC infrastructure across Kubernetes, cloud, and bare-metal environments. The role requires production operations, infrastructure-as-code, GitOps, observability, incident response, and on-call experience.

About the job

Responsibilities

  • Deploy, operate, and maintain blockchain RPC nodes across multiple chains and geographic regions.
  • Manage Kubernetes clusters underlying the blockchain node platform.
  • Perform rolling upgrades and hard fork migrations for blockchain clients across EVM and non-EVM chains.
  • Participate in on-call rotations, triage live incidents via PagerDuty, and coordinate resolution for node outages, latency spikes, and SLO breaches.
  • Develop and maintain AI agents and automation tooling for health checks, auto-healing, and hard fork notifications.
  • Deploy and manage services using ArgoCD and GitOps workflows with Helm charts.
  • Manage bare-metal and cloud infrastructure, including provisioning, benchmarking, and hardware replacement.
  • Respond promptly to security advisories and coordinate upgrades with minimal downtime.
  • Contribute to postmortems and asynchronous review processes; track action items and follow up on resolutions.
  • Collaborate with product, customer success, and engineering teams on chain deprecations, capacity planning, and SLO reporting.

Requirements

  • Experience designing and operating large-scale, multi-region, multi-cloud production systems.
  • Experience with Kubernetes, including StatefulSets, storage management, Secrets, and service mesh technologies such as Istio.
  • Experience with secrets management and access control in multi-cluster environments.
  • Familiarity with automation frameworks for node health checks, upgrades, and remediation workflows.
  • Experience with infrastructure as code, such as Terraform, Ansible, Pulumi, CloudFormation, Chef, or Puppet.
  • Experience with GitOps tooling, including ArgoCD and Helm.
  • Proficiency with cloud infrastructure and bare-metal management, including storage provisioning and snapshot management.
  • Strong understanding of observability tooling, including Grafana, Prometheus, and Alertmanager, with experience building or tuning dashboards and alert rules.
  • Comfort working in an on-call environment and triaging production incidents using PagerDuty and structured runbooks.
  • Ability to write technical documentation and postmortems and contribute to asynchronous team communication.
  • Experience with networking and configuring or managing VPC networks.
  • Basic understanding of security best practices.
  • Passion for blockchain technologies and Web3.

Nice-to-Haves

  • Experience with service mesh deployments such as Istio.
  • Good understanding of web applications and microservice architecture.
  • Experience working with startups.

Compensation and Benefits

  • Attractive salary package.
  • Private medical insurance.
  • Flexible time away.
  • Internal off-site hackathons.
  • Access to a company-rented hacker house during summer.
  • Opportunities to travel across offices.

Skills

Kubernetes, Terraform, Ansible, Pulumi, CloudFormation, Argo CD, Helm, Istio, Grafana, Prometheus, Alertmanager, Pagerduty, GitOps, Vpc Networking, Infrastructure As Code

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Phantom

Phantom

Remote

Staff DevOps Engineer
No salary listedRemote8+ YOEDevOps / SRE

Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Kraken

Kraken

United Arab Emirates
Senior Database Administrator - Core Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.