Senior Site Reliability Engineer responsible for monitoring, incident response, and optimizing the reliability, scalability, and performance of Illumio's AWS and Azure cloud infrastructure and SaaS services. Requires 5+ years SRE experience with strong cloud platform expertise.
170k – 196k/yr
On-site5+ YOEDevOps / SRE
About the role
Your Impact
Monitor system performance, application health, and infrastructure metrics using monitoring and logging services, and implement proactive measures to optimize performance and availability.
Oncall duty for production uptime and support for customer escalations.
Release upgrades and maintenance activities including hotfixes and infrastructure updates.
Lead incident response and resolution efforts, conducting root cause analysis, implementing corrective actions, and documenting post-incident reviews.
Implement security best practices and controls in the cloud environments to protect data, applications, and infrastructure, and ensure compliance with regulatory requirements.
Drive continuous improvement initiatives to enhance reliability, scalability, and efficiency of infrastructure and services, leveraging automation and emerging technologies.
Your Toolkit
Bachelor’s degree in computer science, Engineering, or related field; or equivalent work experience.
5+ years of experience working as a Site Reliability Engineer (SRE) or similar role, with a focus on AWS and/or Azure cloud platform.
Hands-on experience in designing, deploying, and managing AWS and/or Azure infrastructure, including compute, storage, networking, and security services.
Proficiency in scripting and programming languages such as PowerShell, Python, or Go for automation and infrastructure management tasks.
Strong understanding of CI/CD principles and experience with tools such as Azure DevOps, Jenkins, or GitLab CI/CD.
Experience with containerization technologies (e.g., Docker, Kubernetes) and microservices architecture in AWS and Azure environments is a plus.
Excellent analytical, problem-solving, and communication skills, with the ability to collaborate effectively with cross-functional teams.
AWS or Azure certifications such as AWS/Azure Solutions Architect, Azure DevOps Engineer, or Azure Security Engineer are preferred.
Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.
170k – 205k/yr
On-site5+ YOEDevOps / SRE
Senior Network Engineer
Lightning AINew York, NY +2
Senior Network Engineer responsible for designing, deploying, and optimizing large-scale NVIDIA InfiniBand fabrics and UFM for AI/ML GPU clusters. Requires 10+ years data center networking experience with deep expertise in InfiniBand, spine-leaf architectures, automation, and HPC environments.
170k – 210k/yr
On-site10+ YOEDevOps / SRE
Software Engineer, Compute Infrastructure
RenderSan Francisco, CA
Build and own core compute infrastructure for Render's cloud platform, including Kubernetes clusters on hyperscalers and bare metal. Design, scale, debug, and optimize large-scale orchestration, scheduling, and distributed systems with deep Kubernetes and systems expertise.
Senior engineer focused on developer productivity and cloud infrastructure. Designs scalable internal tools, re-architects build systems, and improves CI/CD workflows using Terraform, Go/Python/C++.
170k – 240k/yr
Hybrid5+ YOEDevOps / SRE
Senior Software Engineer - Observability and Reliability
SigmaSan Francisco, CA +1
Build observability tools and platforms (metrics, logging, tracing, alerting) using Go, OpenTelemetry, and Kubernetes. Requires 5+ years experience building high-quality software that other engineers use.