Skip to content
F2F2

Staff Software Engineer, Infrastructure

Hands-on Infrastructure Tech Lead building and scaling AWS cloud infrastructure from scratch for an AI-driven enterprise analytics platform. Owns architecture, IaC, security/compliance (SOC 2), and operational excellence.

About the job

What you'll do

  • Hands-on platform building. Architect and implement foundational cloud infrastructure from scratch: compute, networking, CI/CD, and observability. Actively writing infrastructure-as-code and shipping production systems.
  • Own the infrastructure architecture. Define and execute the technical vision and multi-year roadmap for our AWS-based platform (ECS/Fargate, containerized Python and Node services).
  • Run durable, AI-heavy workloads. Operate and scale our workflow orchestration layer (Temporal), streaming pipelines, vector search infrastructure, and high-throughput LLM inference paths.
  • Design for security and compliance. Build infrastructure that meets SOC 2 requirements from day one: multi-tenant isolation, secrets management, least-privilege IAM, audit logging, and encrypted data flows.
  • Establish operational excellence. Set and uphold standards for IaC, deployment pipelines, incident response, SLOs, and on-call practices; mentor engineers.
  • Cross-functional collaboration. Partner with product, backend, frontend, and enterprise customers to translate requirements into pragmatic infrastructure solutions.

What you bring

  • 7+ years building and scaling mission-critical cloud infrastructure on AWS and/or GCP, with demonstrated experience architecting platforms from the ground up.
  • Production experience with ECS/Fargate, Kubernetes, or equivalent: including service networking, autoscaling, zero-downtime deploys, and multi-environment release strategies.
  • Strong command of Terraform, Pulumi, or CloudFormation, plus CI/CD pipeline design (GitHub Actions or similar) and GitOps workflows.
  • Familiarity operating Postgres at scale (RDS, Supabase, or self-managed), Redis, message/workflow systems (Temporal, SQS, Kafka), and ideally vector databases or LLM serving infrastructure.
  • Practical experience with SOC 2 (or similar) compliance programs, IAM design, VPC architecture, secrets management, and multi-tenant data isolation.
  • Execution-driven mindset with end-to-end ownership of systems.

Skills

AWS, GCP, ECS, Fargate, Kubernetes, Terraform, Pulumi, CloudFormation, GitHub Actions, GitOps, Postgres, Redis, Temporal, SQS, Kafka

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Operations
$195k+/yrOn-site8+ YOEDevOps / SRE

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.

Shield AI

Shield AI

San Diego, CA

Senior Staff Lead Site Reliability Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.

Airbnb

Airbnb

United States

Staff Software Engineer, Service Tools
$212k+/yrRemote9+ YOEDevOps / SRE

Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.

Temporal

Temporal

United States

Staff Software Engineer, Traffic
$212k+/yrRemote8+ YOEDevOps / SRE

Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.