Skip to content

Senior Infrastructure Software Engineer

Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.

About the job

Responsibilities

Infrastructure Software & Automation

  • Design, build, and operate production services, APIs, tooling, and automation for large-scale bare-metal and GPU infrastructure.
  • Build software systems for infrastructure provisioning, configuration, monitoring, and lifecycle management.
  • Develop reliable automation that reduces manual operational work and improves consistency and scalability.
  • Build tools and interfaces for programmatic interaction with physical infrastructure.

Provisioning & Lifecycle Management

  • Develop systems for server discovery, provisioning, configuration, validation, and lifecycle management.
  • Automate infrastructure bring-up and capacity deployment across large fleets of compute systems.
  • Integrate software and automation with hardware management and provisioning systems.
  • Improve tooling and workflows across the infrastructure production lifecycle.

Reliability & Observability

  • Build telemetry, logging, and observability capabilities for infrastructure health visibility.
  • Develop tools and automation to identify, diagnose, and respond to infrastructure and hardware issues.
  • Convert recurring operational issues and hardware failure modes into improvements to software, tooling, and automation.
  • Improve the reliability and scalability of infrastructure operations.

Architecture & Collaboration

  • Partner with Network, Infrastructure Operations, Data Center, and Platform Engineering teams to define requirements and deliver infrastructure capabilities.
  • Participate in architectural discussions and help define the technical direction of infrastructure software.
  • Make pragmatic design decisions balancing reliability, scalability, simplicity, and execution speed.
  • Write design documents and technical documentation and contribute to engineering best practices.

Requirements

  • 8+ years of professional software engineering, infrastructure engineering, or related experience.
  • Strong software engineering fundamentals and experience building production backend systems in Python or similar object-oriented languages.
  • Strong experience with Linux in production environments.
  • Experience building APIs, tooling, or automation for managing infrastructure at scale.
  • Familiarity with containerization and orchestration concepts.
  • Understanding of HPC and bare-metal infrastructure fundamentals, including provisioning and out-of-band management.
  • Ability to navigate ambiguity and make pragmatic architecture decisions for long-term reliability and scale.
  • Experience in a startup or other fast-paced environment with substantial ownership and autonomy.

Nice-to-Haves

  • Experience with bare-metal hardware troubleshooting and provisioning, including PXE/iPXE, BMC, Redfish, or IPMI, particularly with Dell hardware.
  • Experience with GPU servers in bare-metal or virtualized environments.
  • Experience with network switches, routers, and firewalls, particularly SONiC switches, Palo Alto firewalls, or Juniper Networks.
  • Experience with high-performance storage systems, particularly VAST.
  • Experience supporting AI/ML or HPC infrastructure at scale.

Compensation & Benefits

  • Annual base salary: $180,000–$220,000 USD.
  • Discretionary bonus and meaningful equity component.
  • Comprehensive medical, dental, and vision coverage.
  • 401(k) matching in the U.S. and pension contributions in the U.K.
  • Unlimited PTO, company holidays, and floating holidays.
  • Two-week company-wide winter break.
  • Paid parental and family leave.
  • Annual learning and development allowance.
  • Wellness and work-from-home stipends.
  • Four-week paid sabbatical after four years of service.
  • Flexible schedules and hybrid work model.
  • Complimentary in-office meals.

Skills

Python, Linux, APIs, Infrastructure Automation, Bare-Metal Infrastructure, Gpu Infrastructure, Hpc, Containerization, Orchestration, Pxe/Ipxe, Redfish, Ipmi, Kubernetes, Observability, Sonic

Runpod

Runpod

United States

Senior HPC Storage Engineer
$180k+/yrRemote8+ YOEDevOps / SRE

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.

Onebrief

Onebrief

Colorado Springs, CO

Senior Site Reliability Engineer, Colorado Springs
$180k+/yrOn-site5+ YOEDevOps / SRE

Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.