Staff Infrastructure Software Engineer owning multi-cloud (AWS/GCP) platform, Kubernetes/Istio, Terraform IaC, networking to edge printers, and CI/CD pipelines at Carbon. Requires 7+ years production cloud infrastructure experience, expert Terraform, strong Kubernetes and networking skills.
210k – 314k/yr
On-site7+ YOEDevOps / SRE
About the role
Responsibilities
Operate and extend multi-cloud infrastructure (AWS and GCP) for Tier-0 services, ensuring resilience, cost-efficiency, and security.
Design and maintain networking between cloud services and a global fleet of printers, including VPCs, peering, DNS, load balancing, TLS, and secure edge connectivity.
Own and improve Kubernetes clusters and Istio service mesh: cluster architecture, workload isolation, traffic management, mTLS, and observability.
Take end-to-end ownership of Terraform initiatives, including module design, state management, and IaC standards.
Lead improvements to CI/CD pipelines in Jenkins and GitHub Actions for faster builds, safer deploys, and better feedback loops.
Participate in on-call, drive incident response, postmortems, and help teams adopt strong operational practices.
Provide technical leadership, design reviews, code reviews, and mentorship across infrastructure, security, and application teams.
Requirements
7+ years of experience building and operating production cloud infrastructure, with deep expertise in at least one of AWS or GCP and working proficiency in the other.
Strong networking fundamentals: VPCs, subnets, routing, DNS, load balancing, TLS, VPNs, and secure connectivity between cloud and on-prem or edge devices.
Expert-level Terraform: designing modules, managing state at scale, leading migrations without downtime.
Hands-on experience building and maintaining CI/CD pipelines in Jenkins and GitHub Actions, including build performance, artifact management, and safe deployment patterns.
Strong Kubernetes expertise: operating production clusters at scale, control plane and networking model, workload/upgrade/multi-tenant strategies.
Hands-on experience with a service mesh in production (ideally Istio), including envoy-based traffic management, mTLS, authorization policy, and mesh observability.
Fluency in Linux, containers, and at least one scripting or systems language (Python, Go, Bash, or similar).
Production mindset focused on failure modes, blast radius, cost, and observability; experience carrying a pager for infrastructure.
Strong communication, debugging, and problem-solving skills; track record of driving major initiatives end-to-end.
Bachelor's degree in Computer Science or Engineering, or equivalent practical experience.
Nice-to-Haves
Experience connecting cloud services to physical devices or edge hardware in the field.
Familiarity with Bazel or another polyglot monorepo build system.
Experience with security and compliance controls (SOC 2, ISO 27001) for cloud infrastructure.
Experience running or contributing to a platform team serving other engineers as internal customers.
Lead the design and build of internal developer platforms and AI guardrails at Zocdoc. Focus on enabling both engineers and non-technical teams with secure, scalable, easy-to-use tools, CI/CD, and AI workflows while driving adoption through empathy and measurable outcomes. Requires 7+ years platform experience and a Bachelor's degree.
210k – 270k/yrHybrid7+ YOEDevOps / SRE
Staff Software Engineer, App Platform
FieldguideSan Francisco, CA
Lead design and evolution of core platform services, APIs, and shared primitives that power every product surface and AI agent. Drive technical standards and architecture across SaaS, enterprise, and government environments while mentoring engineers.
210k – 265k/yrRemote10+ YOEDevOps / SRE
Staff Software Engineer, Systems Engineering Focus
CrusoeSan Francisco, CA
Designs, builds, and scales customer-facing managed services with a focus on edge agents running on customer infrastructure. Provides technical oversight for high-reliability systems using eBPF, Kubernetes, and low-level Linux metrics; leads cross-team collaboration and mentors engineers.
Build and lead full-stack systems for real-time monitoring and automation of autonomous robotaxi fleets. Requires 8+ years experience, strong JavaScript/TypeScript/Node.js skills, and front-end framework expertise.
210k – 305k/yrHybrid8+ YOEDevOps / SRE
Staff Production Engineer, Compute
CrusoeSan Francisco, CA
Staff Production Engineer responsible for developing automation/observability, scaling virtualization (KVM/QEMU), optimizing Linux kernel performance, and supporting AI/HPC workloads on CPU/GPU/DPU hardware. Requires 8+ years in Linux systems engineering, kernel internals, and virtualization.