Senior Site Reliability Engineer building agentic AI platforms, MCP servers, and automation tools to achieve zero KTLO. Drives SecDevOps culture with focus on security, reliability, observability while partnering on product roadmap and incident response. Requires 6+ years SRE/DevOps experience plus deep expertise in AWS, EKS, Terraform, Kubernetes.
170k – 190k/yr
Hybrid6+ YOEDevOps / SRE
About the role
What you will do
Partner with a team of high-performing engineers and developers who are focused on delivering best in class software products
Function as a Change Agent to introduce and evangelize our shift to a SecDevOps culture, solving for security, reliability, cost-effectiveness, and observability
Building Zero trust security designs and adapting DevSecOps mindset when building new services
Cultivate an automation-first attitude and work to champion code-centric solutions throughout our department to improve velocity and deliverability
Participate in the design process, representing, solving and planning for security, reliability and cost effectiveness prism in the product roadmap
Collaborate and communicate with cross-functional colleagues belonging to the same job family to drive SecDevOps centric cross-org initiatives, tooling and standards
Participate in incident response in collaboration with application owners and the platform team
Build tools that empower capacity planning and demand forecasting, software performance analysis, and system tuning
Design and build agentic AI platforms and AI agents — including home-grown MCP (Model Context Protocol) servers — to drive our organization toward a 0 KTLO north star, shifting engineering effort away from routine operational maintenance and toward high-value, forward-looking initiatives
Develop tooling that enable our product teams to through self service
Think implementing long-term solutions that are complete mechanisms by building tools, driving adoption and inspecting results for tuning
What you bring to the role
At least 6+ years of experience as SRE, DevOps or equivalent engineering roles
Hands-on experience building AI SRE/DevOps Agents and using AI agentic frameworks
Hands-on experience building internal MCPs for AI Agentic use
Strong hands-on experience working with foundational AWS services
Expert level experience working with AWS EKS
Strong hands-on hands-on experience with ArgoCD
Strong hands-on experience Terraform
Experience with at least one language for automation - Bash, Python, Golang, etc.
Experience deploying Serverless application using SAM
Experience building and managing CI/CD pipelines
Strong hands-on experience in Linux architecture, microservices and container orchestration (Docker, Kubernetes, etc.)
Experience with monitoring tools (New Relic, Prometheus, Grafana, Loki etc.)
Experience working on infrastructure projects in an Agile environment
Great communication, collaboration and presentation skills
Ability to team up with people from different disciplines and drive for a win-win
Ability to adapt and build security orchestration and automation at scale
Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.
170k – 205k/yr
On-site5+ YOEDevOps / SRE
Senior Network Engineer
Lightning AINew York, NY +2
Senior Network Engineer responsible for designing, deploying, and optimizing large-scale NVIDIA InfiniBand fabrics and UFM for AI/ML GPU clusters. Requires 10+ years data center networking experience with deep expertise in InfiniBand, spine-leaf architectures, automation, and HPC environments.
170k – 210k/yr
On-site10+ YOEDevOps / SRE
Sr. Site Reliability Engineer
IllumioSunnyvale, CA
Senior Site Reliability Engineer responsible for monitoring, incident response, and optimizing the reliability, scalability, and performance of Illumio's AWS and Azure cloud infrastructure and SaaS services. Requires 5+ years SRE experience with strong cloud platform expertise.
170k – 196k/yr
On-site5+ YOEDevOps / SRE
Software Engineer, Compute Infrastructure
RenderSan Francisco, CA
Build and own core compute infrastructure for Render's cloud platform, including Kubernetes clusters on hyperscalers and bare metal. Design, scale, debug, and optimize large-scale orchestration, scheduling, and distributed systems with deep Kubernetes and systems expertise.