Skip to content
Chess.comChess.com

Site Reliability Engineer

Design and operate multi-regional infrastructure for a high-traffic global gaming platform, owning on-call, monitoring, automation, and hybrid cloud migration to ensure reliability at massive scale.

About the job

Responsibilities

  • Design and implement multi-regional resilient infrastructure capable of handling millions of concurrent sessions and transactions daily across global data centers
  • Lead the hybrid cloud migration strategy, integrating bare-metal datacenter resources with cloud services for optimal performance and cost efficiency
  • Own the on-call rotation and incident response procedures, ensuring rapid resolution of critical system issues and maintaining high availability SLAs
  • Architect monitoring and alerting systems using industry-standard tools to proactively identify and resolve performance bottlenecks before they impact users
  • Collaborate with development teams to implement infrastructure-as-code practices and establish deployment pipelines that support continuous integration and delivery
  • Optimize system performance through capacity planning, load testing, and resource allocation across distributed computing environments
  • Establish and maintain security protocols and risk assessment procedures for infrastructure components and data protection
  • Partner with engineering teams to design scalable solutions for high-traffic applications and real-time processing requirements
  • Drive automation initiatives to reduce manual operational overhead and improve system reliability through scripting and configuration management
  • Mentor team members on SRE best practices and contribute to the development of infrastructure standards and documentation

Requirements

  • Bachelor's degree in Computer Science, Engineering, or related technical field, or equivalent practical experience
  • 5+ years of experience in site reliability engineering, DevOps, or infrastructure engineering roles
  • Strong proficiency with UNIX/Linux operating systems and command-line administration
  • Experience with cloud platforms (GCP, AWS, or Azure) and infrastructure-as-code tools (Terraform, CloudFormation, or similar)
  • Hands-on experience with configuration management systems (Ansible, Chef, Puppet, or similar)
  • Solid understanding of networking fundamentals, protocols (TCP/IP, HTTP/HTTPS, DNS), and network troubleshooting
  • Experience with containerization and orchestration technologies (Docker, Kubernetes, or similar)
  • Proficiency with monitoring and observability tools (Datadog, Prometheus, Grafana, ELK stack, or similar)
  • Experience with relational and NoSQL databases, including performance optimization and scaling strategies
  • Strong collaboration and communication skills for working effectively in a distributed team environment
  • Demonstrated sense of ownership and accountability for system reliability and performance

Nice to Have

  • Experience managing bare-metal server infrastructure and datacenter operations
  • Advanced knowledge of content delivery networks (CDNs) and edge computing
  • Experience with server-side automation and scripting languages (Python, Go, Bash, or similar)
  • Background in high-availability architectures and disaster recovery planning
  • Familiarity with security frameworks and compliance requirements
  • Experience with game server infrastructure or real-time application hosting
  • Knowledge of database administration and optimization for high-concurrency applications
  • Understanding of CI/CD pipelines and deployment automation
  • Experience with capacity planning and performance testing tools
  • Previous experience in a fully remote, distributed work environment
  • Continuous learning mindset with interest in emerging infrastructure technologies

Skills

Linux, GCP, AWS, Azure, Terraform, Ansible, Docker, Kubernetes, Prometheus, Grafana, Datadog, Python, Go, Bash, TCP/IP

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.