Software Engineer, Infrastructure
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
About the job
Responsibilities
- Own the platform foundation, including Google Cloud, Kubernetes, Temporal, GPU infrastructure, deployment, and rollback systems.
- Participate in the on-call rotation and improve incident response, reliability, and actionability.
- Build and operate AI infrastructure for GPU capacity, model training and inference pipelines, and production serving.
- Manage infrastructure cost modeling, metering, attribution, and cost-performance trade-offs.
- Own security boundaries including identity and access management, secrets, least privilege, and software supply-chain integrity.
- Develop infrastructure-as-code, runbooks, observability, tests, standards, and release practices.
- Partner with engineering teams to understand needs, set the platform roadmap, and make architecture decisions.
- Provide architectural direction, code reviews, mentoring, and clear technical documentation.
Requirements
- 8+ years building and operating production distributed systems or server-side systems with a strong infrastructure focus.
- Experience operating production systems where reliability and failure consequences matter.
- Experience carrying a pager, commanding incidents, and performing rollbacks.
- Practical use of SLOs and error budgets.
- Production experience with a major cloud provider, Kubernetes, and infrastructure-as-code.
- Experience owning an architecture or migration from planning through launch.
- Ability to independently identify, scope, gain support for, and deliver unowned work.
- Strong debugging skills, including creating minimal reproductions, analyzing logs, and writing targeted checks.
- Ability to use AI agents effectively and critically in engineering work.
Nice-to-haves
- GPU and machine-learning infrastructure, including capacity planning, training or inference pipelines, and model-serving cost and latency optimization.
- Production security engineering, including IAM, secrets management, software supply-chain security, and least privilege.
- Cloud cost modeling, commitment strategies, reservations, and unit economics.
- CI/CD at monorepo scale and developer-environment tooling.
- Media, video, or GPU-backed workload experience.
- Experience on small teams responsible for broad infrastructure surfaces.
Compensation and Benefits
- Base salary: $220,000–$292,000, plus equity and benefits.
- Healthcare package, 401(k) matching, catered lunches, and flexible vacation time.
Skills
GCP, Kubernetes, Temporal, Gpu Infrastructure, Machine Learning Infrastructure, CI/CD, Infrastructure As Code, Distributed Systems, SLOs, Error Budgets, IAM, Secrets Management, Observability, Monorepos, Cloud Cost Modeling
Similar jobs
DevOps / SRE jobsSenior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.
Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.