Site Reliability Engineer modernizing a multi-cloud (AWS/Azure/GCP) environment into a scalable, observable Kubernetes-based platform using DevOps/SRE practices, AIOps, IaC, and AI-driven automation to support scientific and clinical research programs. Requires 6+ years SRE/DevOps experience with strong Linux, IaC, observability, and scripting skills.
140k – 155k/yr
On-site6+ YOEDevOps / SRE
About the role
Responsibilities
Design and implement enterprise-grade monitoring and observability frameworks (metrics, logs, traces) across distributed systems using enterprise Splunk, Grafana and OpenTelemetry tools.
Establish and manage SLIs, SLOs, and error budgets to drive reliability improvements.
Develop and maintain real-time asset inventory systems across cloud, on-prem, and hybrid environments.
Automate workload onboarding and offboarding processes, ensuring standardization and governance.
Track system ownership, dependencies, and lifecycle states for operational transparency.
Build proactive detection mechanisms using AIOps and intelligent alerting to minimize incident impact.
Design and operate scalable, resilient, and secure infrastructure platforms across cloud and hybrid environments.
Implement automated compliance tracking and enforcement aligned with organizational and regulatory standards (e.g., NIST, FISMA, FedRAMP).
Build and maintain automated deployment environments and pipelines that enforce security, compliance, and operational standards.
Develop “golden paths” and standardized platform templates for consistent workload deployment.
Automate provisioning, patching, configuration management, and environment lifecycle.
Leverage AI/ML coding assistants and vibe coding practices to rapidly develop automation scripts, tools, and internal platforms.
Integrate AI-driven tooling into DevOps pipelines for code quality, security scanning, and operational insights.
Lead adoption of AI-enhanced SRE practices, including intelligent remediation and predictive operations.
Champion DevOps and SRE practices including Infrastructure as Code, CI/CD, observability, and reliability engineering.
Build developer-friendly platforms (“golden paths”) that simplify deployments, reduce friction, and improve velocity.
Enable and optimize infrastructure for AI/ML workloads, including data pipelines, storage systems, and inference environments, GPU-enabled and high-performance compute workloads.
Build and manage containerized and orchestrated platforms (Docker, Kubernetes).
Support cloud migration, modernization, and platform standardization initiatives.
Ensure systems meet security, compliance, backup, and disaster recovery requirements.
Evangelize and promote best practices in DevOps, SRE, and platform engineering to developer communities.
Stay abreast of new technologies in areas including AIOps, MLOps, cloud computing & deployment, site reliability engineering, infrastructure automation, security best practices, data engineering etc.
Requirements
6+ years experience in DevOps / SRE roles with monitoring and observability tools (Prometheus, Grafana, ELK, or cloud-native equivalents) for on-prem and cloud hosted workloads.
4+ years of hands-on Linux experience that includes Ubuntu/CentOS/Red Hat operating systems, containers, dependency management and administration support.
4+ years of experience automating Infrastructure-as-Code (IaC) deployments to Amazon AWS, Google GCP or Microsoft Azure.
4+ years with CI/CD and automation tools such as Terraform, Ansible, Chef, Puppet, Jenkins, GitHub Actions.
Strong scripting skills (Python, Bash, PowerShell or similar).
Proficiency using vibe coding and coding assistants to develop scripts, tools and applications for the DevOps and SRE use cases.
Proficiency to debug or troubleshoot and/or deploying SQL and/or NoSQL databases, object storage, web servers, open-source programming stack for Node.JS, R, Python, .NET Core, Java is desired but not mandatory.
Willingness to learn new technologies, adopt and adapt to emerging technologies or needs from a project to a project.
Cloud certifications preferred.
Certifications in Grafana, Splunk, Docker, Kubernetes preferred but optional.
Nice-to-Haves
Experience optimizing infrastructure for AI/ML workloads.
Familiarity with zero-trust principles and multi-cloud environments (AWS, Azure, GCP).
Build and operate production-critical GitOps deployment platforms, shared service tooling, and infrastructure automation in Go and TypeScript. The role requires 8+ years of software engineering experience plus expertise with Argo CD, Helm, Kubernetes, cloud platforms, and scalable APIs.
140k – 220k/yrRemote8+ YOEDevOps / SRE
DevOps Engineer
Pump.coSan Francisco, CA
Hands-on DevOps role owning AWS infrastructure, building developer tooling, and driving technical roadmap at an early-stage YC startup. Requires 6+ years infra/DevOps experience and strong AWS/K8s/Terraform skills.
140k – 200k/yrOn-site6+ YOEDevOps / SRE
Senior Network Systems Engineer
ForterraEast Palo Alto, CA +2
Deploys, operates, and troubleshoots network infrastructure including routers, switches, Linux appliances, and AWS resources for edge-deployed communications in DDIL environments. Requires 5+ years network engineering experience, Linux proficiency, IaC automation, and 50% domestic travel.
140k – 185k/yrHybrid5+ YOEDevOps / SRE
Senior Infrastructure Engineer
ScrunchNew York, NY +16
Senior Infrastructure Engineer designs, builds, and operates cloud infrastructure, developer tooling, observability, and reliability systems at scale, primarily on GCP. Requires high-velocity dev experience, IaC, database scaling, workflow orchestration, and production Python coding.
140k – 200k/yrRemoteDevOps / SRE
Senior Infrastructure Engineer - Postgres
ClickhouseUnited States
Senior Infrastructure Engineer owns reliability, operations, and automation for ClickHouse's Postgres integration across multi-cloud environments. Requires 7+ years SRE/DevOps experience, Postgres expertise, Terraform/Kubernetes proficiency, and strong Go skills.