Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.
Salary not listed
Hybrid6+ YOEDevOps / SRE
About the role
Responsibilities
Deploy, configure, and debug container-based services using Docker.
Build and own software solutions for cluster operations, including monitoring platforms, workflow automation systems, operational dashboards, and reliability tooling.
Collaborate with cross-functional teams to translate operational requirements into scalable operations and maintenance products and platform capabilities.
Develop APIs, automation services, and integrations that improve operational visibility, incident response, and fleet management across global AI infrastructure.
Manage and operate multiple advanced AI compute infrastructure clusters.
Monitor and oversee cluster health, proactively identifying and resolving potential issues.
Maximize compute capacity through optimization and efficient resource allocation.
Provide 24/7 monitoring and support using automated tools and hands-on troubleshooting.
Handle engineering escalations and collaborate with other teams to resolve complex technical challenges.
Stay current with advancements in AI compute infrastructure and related technologies.
Requirements
6–8 years of relevant experience managing and operating complex compute infrastructure, preferably in machine learning or high-performance computing.
Proficiency in Python and Go, including experience building operational platforms, workflow automation systems, and reliability tooling for large-scale infrastructure environments.
Expertise in distributed systems.
Deep understanding of Linux-based compute systems and command-line tools.
Extensive knowledge of Docker containers and orchestration platforms such as Kubernetes.
Ability to troubleshoot and resolve complex technical issues efficiently.
Experience with monitoring and alerting systems.
Proven ability to own challenges and drive them to completion.
Strong communication and collaboration skills.
Ability to work effectively in a fast-paced environment.
Willingness to participate in a 24/7 on-call rotation.
Preferred Skills
Experience operating and managing large-scale AI clusters.
Knowledge of Ethernet, RoCE, and TCP/IP.
Knowledge of cloud computing platforms such as AWS, Google Cloud, and Azure.
Benefits
Opportunity to build and operate a breakthrough AI platform beyond GPU constraints.
Opportunities to publish and open-source AI research.
Work on one of the world's fastest AI supercomputers.
Startup vitality with job stability.
A non-corporate work culture that respects individual beliefs.
Skills
PythonGoDistributed SystemsLinuxDockerKubernetesmonitoring and alertingethernetroceTCP/IPAWSGCPAzure
Senior Software Engineer - Snowpark Container Service
SnowflakeBellevue, WA +1
Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.
200k – 288k/yrHybrid7+ YOEDevOps / SRE
Senior Software Engineer, Infrastructure & Systems
AstronomerNew York, NY
Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.
200k – 300k/yrHybrid5+ YOEDevOps / SRE
Senior Manager, Site Reliability Engineering - Infrastructure Platform
OktaBellevue, WA
Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.
Build and operate scalable, fault-tolerant cloud infrastructure while leading efficiency initiatives across compute, storage, and networking. The role requires 5+ years of distributed-systems software development experience and expertise with public cloud, infrastructure as code, and cloud-native technologies.
Build and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.