Staff Software Engineer (SRE)
Staff SRE manages and scales Kubernetes clusters on AWS EKS, automates infrastructure with IaC tools, optimizes performance, maintains blockchain nodes and databases, and improves system reliability using monitoring tools. Requires 5+ years SRE experience with strong Kubernetes and AWS proficiency.
About the job
Responsibilities
- Kubernetes Ownership: Manage and scale Kubernetes clusters on AWS EKS, ensuring reliability, performance, and security.
- Infrastructure Automation: Implement and maintain Infrastructure-as-Code (Terraform/Pulumi) to automate infrastructure provisioning and management.
- Performance Optimization: Monitor and optimize system performance, scalability, and resource utilization.
- Blockchain Infrastructure: Configure and maintain crypto nodes across multiple blockchains to support our wallet’s operations.
- Database Scaling: Optimize and scale database infrastructure to handle terabytes of blockchain data efficiently.
- System Reliability: Continuously improve system uptime, monitoring, and observability using tools like Datadog and OpenTelemetry.
- Collaboration: Work closely with backend and product teams to support feature development and system scaling.
Qualifications
- 5+ years in a SRE or Software Engineer role.
- Strong hands-on experience with Kubernetes (EKS) in production environments.
- Proficiency with AWS infrastructure and services (EC2, S3, RDS, IAM).
- Solid experience with Docker and Infrastructure-as-Code tools like Terraform or Pulumi.
- Monitoring and observability experience using tools like Datadog or OpenTelemetry.
Benefits
- Competitive salary and equity
- Eligible to participate in the Company's performance bonus program
- Comprehensive insurance (medical/dental/vision) — 100% covered
- Stipend for your ideal remote set-up
- Flexible hours and a supportive remote environment
- Unlimited vacation
- 401(k) retirement plan
- Monthly wellness benefit
- Weekly meal benefit
- Global off-sites
Target base salary: $200,000 to $250,000 with equity and benefits.
Skills
Kubernetes, Aws Eks, Terraform, Pulumi, Docker, Datadog, OpenTelemetry, Aws Ec2, Aws S3, Aws Rds
Similar jobs
DevOps / SRE jobsOwn reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.
Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.