Senior Software Engineer, Infrastructure & Systems
Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.
200k – 300k/yr
Hybrid5+ YOEDevOps / SRE
About the role
Responsibilities
Design and own the architecture of systems that stand up, scale, configure, and manage the infrastructure running Airflow.
Lead infrastructure efforts end-to-end, including design, implementation, testing, documentation, and rollout.
Design and evolve REST and gRPC APIs for infrastructure lifecycle management, including data modeling, versioning, and backward compatibility.
Own networking and security posture, including authentication, authorization, secure communication, CVE remediation, image hardening, and Pod Security Standards compliance.
Architect observability and traceability across service and network boundaries using metrics, logs, and distributed traces.
Participate in the on-call rotation, diagnose and resolve incidents, contribute to post-mortems, and track remediation work.
Mentor engineers and improve design, code review, testing, and operational practices.
Requirements
5+ years of experience in infrastructure, platform, or systems engineering, with a history of owning production systems at scale.
Deep experience with Kubernetes, including custom Operators, CRDs, and Helm-based deployments.
Experience designing REST or gRPC APIs for customer-facing interfaces or programmatic integrations, including versioning and backward compatibility.
Strong fundamentals in networking, distributed systems, reliability engineering, and security.
Proficiency in Go and/or TypeScript, or the ability to learn new programming languages quickly.
Practical experience designing or building observability and distributed tracing for multi-layer systems.
Ability to lead technical design across teams and communicate trade-offs to engineers and stakeholders.
Experience mentoring engineers and improving team-wide practices.
Nice to Have
Experience designing or operating multi-tenant control planes.
Experience shipping software into air-gapped or highly regulated environments.
Familiarity with Apache Airflow internals or other data orchestration platforms.
Incident command or on-call leadership experience.
Compensation and Benefits
Estimated total compensation: $200,000–$300,000, based on leveling and geography.
Equity component and comprehensive benefits package.
Senior Software Engineer - Snowpark Container Service
SnowflakeBellevue, WA +1
Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.
200k – 288k/yrHybrid7+ YOEDevOps / SRE
Lead Site Reliability Engineer
GleanPalo Alto, CA
Leads SRE team to ensure high availability, scalability, and reliability of cloud services through automation, incident management, and technical leadership. Requires 8+ years SRE experience, team management, and expertise in cloud platforms and containerization.
200k – 260k/yrHybrid8+ YOEDevOps / SRE
Senior Software Engineer, Snowpark Container Service
SnowflakeBellevue, WA +1
Senior Software Engineer building Snowflake's Snowpark Container Services, a managed Kubernetes-based platform for running containerized applications inside Snowflake's Data Cloud. Lead engineering efforts on highly scalable, reliable, multi-tenant container compute infrastructure.
200k – 288k/yrHybrid7+ YOEDevOps / SRE
Senior Software Engineer, Cloud Infrastructure
DecagonSan Francisco, CA +1
Build and operate scalable cloud infrastructure platforms, abstractions over Kubernetes and major cloud providers, and secure enterprise deployments for Decagon's agentic AI systems. Requires 4+ years in infrastructure/DevOps with deep Terraform, Kubernetes, and cloud networking experience.
200k – 400k/yrOn-site4+ YOEDevOps / SRE
Senior Software Engineer, Observability
Together AISan Francisco, CA
Senior Software Engineer building scalable observability platforms (metrics, logs, traces) with Prometheus, Grafana, OpenTelemetry and related tools for Together AI's GPU cloud infrastructure. Requires strong distributed systems and infrastructure-as-code experience.