AI Infrastructure Engineer, Sandbox Platform
Build and operate a secure, high-performance sandboxing platform for agentic code execution across containerized and virtualized environments. The role requires systems software experience, deep Linux knowledge, proficiency in Go, Rust, or C/C++, and strong developer-facing API and production debugging skills.
About the job
Responsibilities
- Design and build the sandboxing platform, client library, and API surface for secure code execution across containerized and virtualized environments.
- Ensure strong isolation, security, and reproducibility across user sessions and workloads.
- Optimize cold-start latency, memory footprint, and resource utilization at scale.
- Reduce error rates through systematic debugging, monitoring, and proactive fixes.
- Partner with internal teams to understand platform needs, debug issues, and build supporting tooling.
- Respond to incidents and production issues, conduct root-cause analysis, and implement preventive fixes.
- Help develop and maintain the sandboxing product roadmap, balancing immediate needs with long-term architecture.
- Lead architecture reviews and own projects end-to-end from design through deployment.
Requirements
- 4+ years of experience building high-performance systems software, including meaningful experience maintaining libraries, SDKs, or developer-facing APIs.
- Deep understanding of Linux internals, including process isolation, memory management, cgroups, and namespaces.
- Experience with containerization and virtualization technologies such as Docker, Firecracker, gVisor, QEMU, or Kata Containers.
- Proficiency in a systems programming language such as Go, Rust, or C/C++.
- Strong focus on developer experience, including API design, error propagation, documentation, and library quality.
- Ability to work across infrastructure layers, from kernel modules to orchestration frameworks such as Kubernetes.
- Strong debugging skills and ability to navigate performance and security tradeoffs in production systems.
- Comfort with ambiguity and switching between incident response and proactive product development.
Nice to Have
- Experience as a founder or early infrastructure startup engineer with end-to-end product ownership.
- Familiarity with LLM agents and agent frameworks such as OpenHands, Agent2Agent, or MCP.
- Experience running secure workloads in multi-tenant or untrusted environments, including FaaS, CI sandboxes, or remote notebooks.
- Exposure to snapshotting and restore techniques such as CRIU, VM snapshots, or overlays.
- Open-source contributions to systems or developer-tools projects.
- Production on-call or incident-response experience.
Skills
Linux, Go, Rust, C/C++, Docker, Firecracker, Gvisor, Qemu, Kubernetes, Cgroups, Namespaces, Criu, Virtualization, API Design, Llm Agents
Similar jobs
Backend Engineering jobsBuild and scale Ruby on Rails backend services for rewards, incentives, and loyalty features in a high-throughput platform. The role owns complex feature delivery, contributes to architecture, mentors junior engineers, and requires 3+ years of professional software engineering experience.
Build and operate backend capabilities for a cloud identity platform, including authentication flows, APIs, and developer experiences. The role requires 3+ years of experience with high-scale production systems and RESTful API development, with Go, TypeScript, and DynamoDB as preferred skills.
Build and scale reliable research infrastructure and distributed systems for evolving AI research workflows. The role independently leads complex technical projects, makes foundational architectural decisions, and partners with researchers and engineering teams.
Build and operate Internet-scale HTTP and TLS infrastructure, migrate services to a Rust-based proxy, and improve protocol performance. The role requires systems programming experience, strong reliability and security practices, and interest in open-source standards.
Build and evolve shared platform infrastructure, data systems, and developer tooling for AI products. The role requires 5+ years of backend engineering experience, distributed systems expertise, and familiarity with cloud, orchestration, databases, and CI/CD technologies.