AI Inference Core - Senior SW Engineer for Platform & DevOps
Build and operate the platform layer behind Cerebras engineering infrastructure, including CI/CD, Kubernetes, deployment automation, cloud and on-premises systems, developer environments, and observability. The role requires 5+ years of infrastructure or software engineering experience and strong debugging and systems fundamentals.
Salary not listed
Hybrid5+ YOEDevOps / SRE
About the role
Responsibilities
Design, build, and maintain CI/CD systems supporting build, test, integration, qualification, and release workflows.
Build and operate Kubernetes-based platforms and services used by engineering teams.
Develop deployment systems, internal tools, and self-service workflows that make infrastructure changes repeatable, reviewable, and safe.
Debug issues spanning CI pipelines, Kubernetes workloads, networking, storage, authentication, operating systems, and distributed applications.
Perform root-cause analysis and implement lasting fixes rather than relying on repeated manual intervention.
Partner with software, IT, security, networking, release, and developer-productivity teams to deliver scalable infrastructure solutions.
Requirements
5+ years of professional experience in platform engineering, DevOps, infrastructure engineering, site reliability engineering, or software engineering.
Hands-on experience building or maintaining CI/CD pipelines and automated software-delivery workflows.
Experience deploying and operating services using Kubernetes and containerized environments.
Experience with a major cloud platform, preferably AWS, and programmatic infrastructure provisioning.
Strong understanding of Linux or Unix operating-system fundamentals.
Understanding of networking concepts such as DNS, routing, load balancing, proxies, ports, TLS, and service connectivity.
Proficiency in Python, Shell, or another language used to build infrastructure automation and operational tooling.
Experience with monitoring, logging, alerting, dashboards, and incident investigation.
Strong debugging and problem-solving skills across applications, infrastructure, networking, and operating systems.
Nice-to-haves
Experience with infrastructure-as-code tools, specifically Terraform.
Experience with Kubernetes controllers, operators, custom resources, Helm, Argo CD, or similar platform technologies.
Senior Software Engineer - Snowpark Container Service
SnowflakeBellevue, WA +1
Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.
200k – 288k/yrHybrid7+ YOEDevOps / SRE
Senior Software Engineer, Infrastructure & Systems
AstronomerNew York, NY
Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.
200k – 300k/yrHybrid5+ YOEDevOps / SRE
Cluster Operations Software Engineer
Cerebras SystemsSunnyvale, CA
Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.
Salary not listedHybrid6+ YOEDevOps / SRE
Senior Manager, Site Reliability Engineering - Infrastructure Platform
OktaBellevue, WA
Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.
Build and operate scalable, fault-tolerant cloud infrastructure while leading efficiency initiatives across compute, storage, and networking. The role requires 5+ years of distributed-systems software development experience and expertise with public cloud, infrastructure as code, and cloud-native technologies.