Skip to content
AstronomerAstronomerNew York, NY

Staff Software Engineer, Platform Infrastructure

Staff Software Engineer building foundational multi-cloud platform infrastructure for Astronomer's Astro DataOps platform. Requires deep distributed systems expertise, Kubernetes operator-level knowledge, strong Go proficiency, and experience driving technical strategy at scale.

275k – 377k/yr
Hybrid7+ YOEDevOps / SRE

About the role

What you get to do

  • Own and develop our platform infrastructure strategy, with the sponsorship and responsibility to match. Map out what we need, make the calls, and own the outcomes.
  • Be directly involved in deciding what we work on and how we work on it. Make promises, and keep them.
  • Make principled build vs. buy assessments and advocate for the right tools for the right job — not the fashionable ones, not the ones already in the estate just because they’re there.
  • Create and maintain comprehensive internal documentation and decision records for systems and processes. Participate in architectural forums and make principled, open decisions that the rest of the organisation can learn from and hold us to.

What you bring to the role

  • Distributed systems depth, grounded in practice. You have a solid working model of how production systems fail — consistency and availability tradeoffs, failure cascades, backpressure, graceful degradation. You can draw the diagram, explain the failure modes at each node, and make a reasoned argument for which ones actually matter in a given context. NALSD thinking is how you naturally approach a new system design.
  • Kubernetes at operator depth. You know what happens inside the scheduler and the control loop when things go wrong, because you’ve been there. You’ve operated clusters under real load, not just deployed workloads onto them.
  • Strong Go proficiency. The platform team writes production Go. You should be fluent: you’ve built and shipped systems in it, and you have opinions about what good Go looks like.
  • Multi-cloud experience, not just multi-cloud exposure. You’ve made considered architectural decisions across AWS, GCP, and/or Azure — not just consumed managed services, but evaluated tradeoffs between them and lived with those decisions in production.
  • Experience defining requirements and driving technology choices across an engineering organization. You’ve been the person in the room who frames the decision correctly, not just the one who executes it.
  • Strong written and verbal communication. You can write a design doc that changes minds, and a postmortem that makes the organisation smarter. You’ve worked effectively in a globally-distributed team.

Bonus points if you have

  • Experience with storage primitives at the system level — you’ve reasoned about when to reach for a relational store vs. an object store vs. something else, and you have real opinions informed by real failures.
  • Experience working on a SaaS/PaaS product across multiple cloud providers.
  • Familiarity with Apache Airflow or workflow orchestration systems.

Skills

GoKubernetesAWSGCPAzureDistributed Systemsapache airflow

Similar roles

DevOps / SRE jobs
Wispr Flow

Staff Platform Engineer, Infrastructure

Wispr FlowSan Francisco, CA

Staff Platform Engineer building scalable infrastructure to support low-latency voice AI product used hundreds of times daily by millions. Architect core platform components connecting product and ML, anticipate scaling bottlenecks, improve developer experience, and set technical direction for growing team.

270k – 350k/yrOn-site7+ YOEDevOps / SRE
Temporal

Senior Staff Software Engineer, Infrastructure

TemporalUnited States

Designs and implements large-scale public cloud infrastructure, builds complex distributed systems and microservices. Requires 10+ years experience, expert skills in performance tuning, concurrency, multiple cloud providers like AWS/GCP/Azure, and graduate degree or equivalent.

260k – 325k/yrRemote10+ YOEDevOps / SRE
Postman

Member of Technical Staff, AI Reliability & Monitoring Engineering Lead

PostmanSan Francisco, CA

Lead AI reliability engineering for Postman's API and agentic systems, building monitoring, observability, and automation for high availability. Requires strong SRE/DevOps background in large-scale AI infrastructure and cloud platforms.

256k – 276k/yrHybridDevOps / SRE
Postman

Member of Technical Staff, AI Platform & Architecture (Infrastructure)

PostmanSan Francisco, CA +3

Builds and maintains distributed AI infrastructure for model training, inference, and data pipelines. Requires experience in GenAI systems, distributed computing, Python/Go, and scaling AI workloads on GPUs/cloud.

256k – 276k/yrHybridDevOps / SRE
Earnin

Staff Site Reliability Engineer

EarninMountain View, CA

Lead EarnIn's AI-first reliability engineering strategy. Define SLOs/SLIs, build AI agents for incident response and on-call automation, and partner with engineering teams to embed AI-assisted operations across production systems on AWS.

252k – 308k/yrHybrid7+ YOEDevOps / SRE