Infrastructure Engineer / SRE
Designs and operates large-scale infrastructure for secure, scalable AI agent runtimes, untrusted code execution, and multi-cloud deployments. Requires strong expertise in distributed systems, containers, Kubernetes, and security.
About the job
What you will work on
- Large scale untrusted code execution and sandboxing
- Massively parallel AI agent runtime and scheduling systems
- Multi cloud and customer VPC deployment architecture
Right now we are looking for two main roles:
- Highly reliable distributed systems with strict security and data guarantees
- Deep observability across AI workflows and infrastructure
- Enterprise grade integration and metadata platforms
What we look for
- Strong background in distributed systems, scalability, multi cloud architecture, and security
- Experience with operating systems, containers, K8s, cloud networking, and event driven runtimes like Knative or KEDA
- Passion for building simple, elegant solutions to hard systems problems
- Experience designing, building, and operating large scale infrastructure
- Strong zero to one mindset
Compensation
Base pay range: $175,000 – $275,000 per year.
Skills
Kubernetes, Distributed Systems, Multi-Cloud, Containers, Cloud Networking, Knative, Keda, Sandboxing, Observability, Vpc
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.