Skip to content
SnowflakeSnowflake

Infrastructure Engineer, Observe by Snowflake

Builds and operates scalable AWS infrastructure for Observe by Snowflake's observability platform, focusing on reliability, CI/CD pipelines, and developer tooling. Requires 2+ years in infrastructure/SRE/DevOps with Kubernetes, IaC tools, and cloud experience.

About the job

What You’ll Do

  • Design, build, and operate scalable cloud infrastructure in AWS supporting a high-scale observability platform.
  • Improve system reliability, performance, and operational visibility across development and production environments.
  • Develop and maintain CI/CD pipelines and internal tooling to improve developer productivity and deployment safety.
  • Identify and mitigate security risks, and help maintain internal security standards and compliance requirements.
  • Build infrastructure that supports high availability, scalability, and operational resilience.
  • Participate in an on-call rotation, contributing to incident response and post-incident improvements.
  • Partner closely with engineering teams to ensure infrastructure supports evolving product and platform needs.

What We’re Looking For

  • 2+ years of experience in Infrastructure Engineering, Site Reliability Engineering (SRE), DevOps, or related roles.
  • Experience operating container orchestration platforms such as Kubernetes.
  • Hands-on experience managing cloud infrastructure using Infrastructure-as-Code tools such as Terraform, Ansible, or similar.
  • Strong programming skills in Go, Python, or similar languages, with a focus on automation and systems development.
  • Experience supporting production systems at scale, with a focus on reliability and operational excellence.
  • Strong problem-solving skills and the ability to balance short-term operational needs with long-term infrastructure design.
  • Experience with AWS, GCP, and Azure.

Nice to Have

  • Experience operating large-scale distributed systems.
  • Familiarity with observability platforms, telemetry pipelines, or monitoring infrastructure.
  • Experience improving developer platform tooling or internal infrastructure platforms.
  • Experience working in high-growth or rapidly evolving engineering environments.

Skills

Kubernetes, Terraform, Ansible, AWS, GCP, Azure, Go, Python, CI/CD, Infrastructure As Code

Otter

Otter

Mountain View, CA

Production Engineer
$155k+/yrHybrid2+ YOEDevOps / SRE

Production Engineer builds and operates large-scale systems, focusing on automation, monitoring, infrastructure management, and resilient operations. Requires 2+ years in SRE/DevOps, expertise in Linux, AWS, Kubernetes, and programming in Python or Golang.

Airtable

Airtable

San Francisco, CA
Software Engineer, Infrastructure (2-8 YOE)
$148k+/yrHybrid2+ YOEDevOps / SRE

Backend engineers build and scale Airtable's infrastructure across teams like Base, Compute, Data, Storage, and Traffic. Requires 2-8 years experience in distributed systems, databases; CS degree; hybrid work in SF, NYC, Seattle, or LA areas.

Greptile

Greptile

San Francisco, CA

Infrastructure Engineer
$190k+/yrOn-site1+ YOEDevOps / SRE

Build and operate robust infrastructure, support enterprise deployments, and improve on-premises delivery for a rapidly scaling AI code review platform. The role requires networking expertise, cloud and container experience, and at least one year of infrastructure or software engineering experience.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff, Systems Infrastructure
$200k+/yrOn-siteDevOps / SRE

Build and operate large-scale scheduling, storage, caching, and networking infrastructure for AI training and inference. The role targets PhD researchers graduating by December 2026 with systems research depth and strong programming and performance-measurement skills.

Mercury

Mercury

San Francisco, CA
Software Engineer - Infrastructure
$116k+/yrRemote2+ YOEDevOps / SRE

Build Mercury’s secure, observable infrastructure platform across AWS, networking, containers, and developer tooling. The role requires strong Linux fundamentals, cloud-native experience, technical writing ability, and software development skills, with opportunities to support AI-agent infrastructure.