Skip to content
SalientSalient

Staff Infrastructure Engineer

Staff Infrastructure Engineer architects and owns scalable cloud infrastructure (AWS/GCP) powering AI-driven financial operations, optimizes GPU workloads, drives reliability via SLOs and monitoring, and enhances developer velocity through CI/CD and platform tools. Requires 5+ years experience with distributed systems.

About the job

What You'll Do

  • Lead architectural decisions and technical reviews for infrastructure-critical initiatives.
  • Design, build, and own the cloud infrastructure (AWS/GCP) that runs Salient - from compute and networking to storage and observability.
  • Develop scalable harnesses that enable coding agents to operate reliably without compromising system stability or code quality.
  • Partner closely with the modeling team to optimize the serving and performance of GPU-intensive workloads.
  • Drive reliability and performance across the stack by defining SLOs, building robust monitoring and alerting, and leading incident response and postmortems.
  • Own developer platform investments that materially improve engineering velocity, including CI/CD, deployment tooling, environments, and internal infrastructure abstractions.
  • Establish infrastructure best practices, patterns, and standards as a technical authority across the engineering org.
  • Identify and reduce technical debt across infrastructure systems, with a focus on long-term scalability and operational health.

What You'll Bring

  • 5+ years of software engineering experience, with 2+ years at the senior or staff level in infrastructure/platform roles, working on large-scale distributed systems.
  • Deep expertise in cloud platforms (AWS or GCP) - compute, networking, storage, IAM, and cost optimization.
  • Expert in infrastructure-as-code, with a strong track record of building scalable automation systems.
  • Extensive experience owning and scaling Kubernetes and CI/CD systems in high-throughput, production environments.
  • Track record of building and operating high-availability, high-throughput distributed systems with mature observability practices.
  • Strong technical communication - able to document architecture clearly and influence engineering decisions across teams.

Nice to Have

  • Background in security.
  • Exposure to serving AI/ML workloads.
  • Combination of big tech and startup experience.

Skills

AWS, GCP, Kubernetes, CI/CD, Infrastructure As Code, Distributed Systems, Observability, SLOs, IAM, Cost Optimization

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Operations
$195k+/yrOn-site8+ YOEDevOps / SRE

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.

Shield AI

Shield AI

San Diego, CA

Senior Staff Lead Site Reliability Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.

Airbnb

Airbnb

United States

Staff Software Engineer, Service Tools
$212k+/yrRemote9+ YOEDevOps / SRE

Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.

Temporal

Temporal

United States

Staff Software Engineer, Traffic
$212k+/yrRemote8+ YOEDevOps / SRE

Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.