Skip to content
InstacartInstacartUnited States

Site Reliability Engineer II

Supports the reliability and performance of large-scale systems through monitoring, incident response, automation, troubleshooting, and dependable deployments. The role requires 2–4 years of software engineering experience and familiarity with scripting, system administration, and cloud platforms is preferred.

133k – 169k/yr
Remote2+ YOEDevOps / SRE

About the role

Responsibilities

  • Monitor systems and respond to alerts, escalating issues as needed.
  • Assist in incident management, following established protocols and documenting steps.
  • Maintain and improve process and procedure documentation.
  • Develop and maintain automation scripts and tools.
  • Apply site reliability engineering best practices.
  • Support application and service deployments, ensuring smooth and reliable releases.
  • Collaborate with senior engineers to troubleshoot and resolve technical issues.

Requirements

  • 2–4 years of software engineering experience.
  • Understanding of programming concepts and scripting languages.
  • Strong interest in system administration and troubleshooting.
  • Excellent problem-solving and analytical skills.
  • Strong ownership of incident handling.
  • Ability to work effectively in a team and communicate clearly.
  • Eagerness to learn and grow in site reliability engineering.

Nice to Have

  • Experience with Ruby or Go.
  • Familiarity with cloud platforms such as AWS, Google Cloud, or Azure.
  • Previous site reliability engineering experience.

Compensation and Benefits

  • Market-competitive compensation and benefits based on work location.
  • Eligible for a new-hire equity grant and annual refresh grants.
  • Remote work policy and benefits offerings are available to employees.
  • ATS-listed base salary: $133,000–$169,000 USD.

Skills

RubyGoAWSGCPmicrosoft azureScriptingsystem administrationIncident ManagementMonitoringAutomationDistributed Systems

Similar roles

DevOps / SRE jobs
Chime

Software Engineer, Infrastructure

ChimeSan Francisco, CA

Build and operate foundational data infrastructure at Chime, owning deployment platforms (Airflow, Flink) and core storage (DynamoDB, RDS). Requires 2+ years infrastructure/backend experience, Terraform, Kubernetes, AWS, and Python.

133k – 184k/yrHybrid2+ YOEDevOps / SRE
Chime

Software Engineer, Infrastructure

ChimeUnited States

Build and operate foundational data infrastructure including Airflow, Flink, DynamoDB, and RDS using Terraform and Kubernetes. Requires 2-4 years of infrastructure/platform experience and strong Python skills.

133k – 184k/yrRemote2+ YOEDevOps / SRE
Yext

Systems Engineer

YextNew York, NY

Design, automate, and maintain reliable infrastructure across cloud and colocation environments. Build monitoring, self-service tools, and standards for distributed systems in a Linux-heavy stack.

137k – 164k/yrOn-site2+ YOEDevOps / SRE
Webflow

Associate Infrastructure Engineer

WebflowUnited States

Operates and improves Webflow’s production infrastructure, focusing on Kubernetes reliability, observability, cloud systems, infrastructure as code, and incident response. The role requires 2+ years debugging distributed systems and comfort with application code and on-call operations.

140k – 190k/yrRemote2+ YOEDevOps / SRE
Cloudflare

Network Deployment Engineer

CloudflareAtlanta, GA +3

Deploys and expands global datacenter and physical network infrastructure, coordinating contractors, vendors, installations, and operational processes. Requires at least two years of datacenter or Linux systems administration experience plus networking, configuration management, scripting, and project coordination skills.

126k – 173k/yrHybrid2+ YOEDevOps / SRE