Software Engineer - Infrastructure (Mid - Senior)
Build and scale secure multi-tenant container infrastructure and code sandboxes on AWS/GCP for an applied AI coding platform. Own reliability, observability, and performance for 500k+ containers per month.
About the job
What You’ll Do
- Design and operate secure, multi-tenant container infrastructure with fast startup and smart autoscaling
- Ship cloud deployments (Helm/Terraform) with SSO, network controls, and audit logging
- Drive observability (metrics, traces, logs) with clear SLOs; lead incident response
- Optimize images, scheduling, networking, and cost; build fair-use and rate-limiting controls
What You Bring
- Production Kubernetes and container internals (Docker/containerd)
- Strong networking fundamentals
- Cloud experience (AWS/GCP/Azure) and IaC (Terraform/Helm)
- Monitoring/Logging (Prometheus, Grafana, OpenTelemetry, ELK/Vector)
- Security best practices for containerized, multi-tenant systems
Nice to Have
- gVisor/Kata/Firecracker
- Cilium/eBPF
- GPU scheduling
- Serverless autoscaling (KEDA/Knative/Karpenter)
- Built an AI side project and enjoy tinkering with LLMs
Compensation & Benefits
- Competitive base salary and meaningful equity
- Health & dental insurance
- Gym reimbursement
- Daily team meals
- Commuter benefits
Skills
Kubernetes, Docker, Containerd, Terraform, Helm, AWS, GCP, Prometheus, Grafana, OpenTelemetry
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.
Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.
Build internal developer platforms, reusable services, and automation that improve software delivery, infrastructure self-service, reliability, and developer productivity. The role requires 5+ years of platform, software, infrastructure, DevOps, or SRE experience plus expertise in cloud-native technologies, Kubernetes, CI/CD, and Infrastructure as Code.