Platform Engineer
Builds backend infrastructure and core platform for AI agent cloud, including VM hypervisors, LLM sandboxes, networking, and orchestration. Requires 5+ years in distributed systems and Linux administration for onsite role in San Francisco.
About the job
Responsibilities
- Designing and building the E2B backend and core infrastructure
- Working with VM hypervisors like Firecracker, gVisor, or Linux systems
- Building and optimizing runtimes and sandboxes for LLMs
- Developing networking solutions for secure, isolated environments
- Monitoring resources and optimizing sandbox performance
- Solving general infrastructure challenges at scale
- Working with orchestration technologies like Kubernetes or Nomad
- Collaborating closely with Distributed Systems Engineers
Requirements
- 5+ years building infrastructure, especially distributed systems
- 5+ years of Linux administration - knowledge of Linux fundamentals: bootloader, kernel, package management, networking, storage, namespaces, containers
- Experience building and operating infrastructure at scale
- Excited to work in person from San Francisco on a DevTool product
- Detail-oriented with great taste in design and engineering
- Comfortable working closely with users
- Proactive, not afraid to take ownership of part of the product
- Excited to take projects from 0 → 1 with the support of the team
Benefits
- Full healthcare, vision, and dental insurance
- Unlimited PTO
Skills
Linux, Kubernetes, Firecracker, Gvisor, Nomad, Distributed Systems, Namespaces, Containers, Networking, Sandboxes
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.