Supports the reliability and performance of large-scale systems through monitoring, incident response, automation, troubleshooting, and dependable deployments. The role requires 2–4 years of software engineering experience and familiarity with scripting, system administration, and cloud platforms is preferred.
133k – 169k/yr
Remote2+ YOEDevOps / SRE
About the role
Responsibilities
Monitor systems and respond to alerts, escalating issues as needed.
Assist in incident management, following established protocols and documenting steps.
Maintain and improve process and procedure documentation.
Develop and maintain automation scripts and tools.
Apply site reliability engineering best practices.
Support application and service deployments, ensuring smooth and reliable releases.
Collaborate with senior engineers to troubleshoot and resolve technical issues.
Requirements
2–4 years of software engineering experience.
Understanding of programming concepts and scripting languages.
Strong interest in system administration and troubleshooting.
Excellent problem-solving and analytical skills.
Strong ownership of incident handling.
Ability to work effectively in a team and communicate clearly.
Eagerness to learn and grow in site reliability engineering.
Nice to Have
Experience with Ruby or Go.
Familiarity with cloud platforms such as AWS, Google Cloud, or Azure.
Previous site reliability engineering experience.
Compensation and Benefits
Market-competitive compensation and benefits based on work location.
Eligible for a new-hire equity grant and annual refresh grants.
Remote work policy and benefits offerings are available to employees.
ATS-listed base salary: $133,000–$169,000 USD.
Skills
RubyGoAWSGCPmicrosoft azureScriptingsystem administrationIncident ManagementMonitoringAutomationDistributed Systems
Build and operate foundational data infrastructure at Chime, owning deployment platforms (Airflow, Flink) and core storage (DynamoDB, RDS). Requires 2+ years infrastructure/backend experience, Terraform, Kubernetes, AWS, and Python.
133k – 184k/yrHybrid2+ YOEDevOps / SRE
Software Engineer, Infrastructure
ChimeUnited States
Build and operate foundational data infrastructure including Airflow, Flink, DynamoDB, and RDS using Terraform and Kubernetes. Requires 2-4 years of infrastructure/platform experience and strong Python skills.
133k – 184k/yrRemote2+ YOEDevOps / SRE
Systems Engineer
YextNew York, NY
Design, automate, and maintain reliable infrastructure across cloud and colocation environments. Build monitoring, self-service tools, and standards for distributed systems in a Linux-heavy stack.
137k – 164k/yrOn-site2+ YOEDevOps / SRE
Associate Infrastructure Engineer
WebflowUnited States
Operates and improves Webflow’s production infrastructure, focusing on Kubernetes reliability, observability, cloud systems, infrastructure as code, and incident response. The role requires 2+ years debugging distributed systems and comfort with application code and on-call operations.
140k – 190k/yrRemote2+ YOEDevOps / SRE
Network Deployment Engineer
CloudflareAtlanta, GA +3
Deploys and expands global datacenter and physical network infrastructure, coordinating contractors, vendors, installations, and operational processes. Requires at least two years of datacenter or Linux systems administration experience plus networking, configuration management, scripting, and project coordination skills.