Senior Infrastructure Software Engineer
Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.
About the job
Responsibilities
Infrastructure Software & Automation
- Design, build, and operate production services, APIs, tooling, and automation for large-scale bare-metal and GPU infrastructure.
- Build software systems for infrastructure provisioning, configuration, monitoring, and lifecycle management.
- Develop reliable automation that reduces manual operational work and improves consistency and scalability.
- Build tools and interfaces for programmatic interaction with physical infrastructure.
Provisioning & Lifecycle Management
- Develop systems for server discovery, provisioning, configuration, validation, and lifecycle management.
- Automate infrastructure bring-up and capacity deployment across large fleets of compute systems.
- Integrate software and automation with hardware management and provisioning systems.
- Improve tooling and workflows across the infrastructure production lifecycle.
Reliability & Observability
- Build telemetry, logging, and observability capabilities for infrastructure health visibility.
- Develop tools and automation to identify, diagnose, and respond to infrastructure and hardware issues.
- Convert recurring operational issues and hardware failure modes into improvements to software, tooling, and automation.
- Improve the reliability and scalability of infrastructure operations.
Architecture & Collaboration
- Partner with Network, Infrastructure Operations, Data Center, and Platform Engineering teams to define requirements and deliver infrastructure capabilities.
- Participate in architectural discussions and help define the technical direction of infrastructure software.
- Make pragmatic design decisions balancing reliability, scalability, simplicity, and execution speed.
- Write design documents and technical documentation and contribute to engineering best practices.
Requirements
- 8+ years of professional software engineering, infrastructure engineering, or related experience.
- Strong software engineering fundamentals and experience building production backend systems in Python or similar object-oriented languages.
- Strong experience with Linux in production environments.
- Experience building APIs, tooling, or automation for managing infrastructure at scale.
- Familiarity with containerization and orchestration concepts.
- Understanding of HPC and bare-metal infrastructure fundamentals, including provisioning and out-of-band management.
- Ability to navigate ambiguity and make pragmatic architecture decisions for long-term reliability and scale.
- Experience in a startup or other fast-paced environment with substantial ownership and autonomy.
Nice-to-Haves
- Experience with bare-metal hardware troubleshooting and provisioning, including PXE/iPXE, BMC, Redfish, or IPMI, particularly with Dell hardware.
- Experience with GPU servers in bare-metal or virtualized environments.
- Experience with network switches, routers, and firewalls, particularly SONiC switches, Palo Alto firewalls, or Juniper Networks.
- Experience with high-performance storage systems, particularly VAST.
- Experience supporting AI/ML or HPC infrastructure at scale.
Compensation & Benefits
- Annual base salary: $180,000–$220,000 USD.
- Discretionary bonus and meaningful equity component.
- Comprehensive medical, dental, and vision coverage.
- 401(k) matching in the U.S. and pension contributions in the U.K.
- Unlimited PTO, company holidays, and floating holidays.
- Two-week company-wide winter break.
- Paid parental and family leave.
- Annual learning and development allowance.
- Wellness and work-from-home stipends.
- Four-week paid sabbatical after four years of service.
- Flexible schedules and hybrid work model.
- Complimentary in-office meals.
Skills
Python, Linux, APIs, Infrastructure Automation, Bare-Metal Infrastructure, Gpu Infrastructure, Hpc, Containerization, Orchestration, Pxe/Ipxe, Redfish, Ipmi, Kubernetes, Observability, Sonic
Similar jobs
DevOps / SRE jobsOwn the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.