Senior Infrastructure Engineer
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
About the job
Responsibilities
- Architect, implement, and roll out large-scale infrastructure projects independently.
- Build and scale infrastructure for agent orchestration, agent sandboxing, and hosted MCP services.
- Help define the technical roadmap and speak directly with customers to inform priorities.
- Work with founders to improve infrastructure reliability and scalability.
- Design and maintain Kubernetes clusters, CI/CD pipelines, and GCP cloud infrastructure.
- Ensure high availability, security, observability, and performance across systems.
Requirements
- Hands-on experience with Kubernetes, Docker, Redis, GCP, Terraform, Prometheus, and CI/CD.
- Experience owning infrastructure at scale in a fast-moving environment.
- Strong software engineering fundamentals and ability to own projects from architecture through production.
- Comfort working in a small, high-ownership team.
- Strong communication skills for customer conversations and collaboration with founders.
Nice-to-haves
- Experience at a high-growth startup scaling infrastructure rapidly.
- Exposure to AI/ML infrastructure and model serving.
- Experience with agent orchestration, sandboxing, MCP, or related distributed systems.
Compensation and Benefits
- Salary of $150,000–$300,000 and meaningful equity.
- 24 days of paid time off for San Francisco employees and 20 days globally.
- $350/month wellness benefit.
- 15 days of temporary remote-work flexibility per year.
- 401(k) match up to 6%.
Skills
Kubernetes, Docker, Redis, GCP, Terraform, Prometheus, CI/CD, Distributed Systems, Ai/Ml Infrastructure, Model Serving
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.
Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.
Build internal developer platforms, reusable services, and automation that improve software delivery, infrastructure self-service, reliability, and developer productivity. The role requires 5+ years of platform, software, infrastructure, DevOps, or SRE experience plus expertise in cloud-native technologies, Kubernetes, CI/CD, and Infrastructure as Code.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.