# Senior Infrastructure Software Engineer

**Company:** [Lightning AI](https://hotfix.jobs/companies/lightning-ai)
**Location:** New York, NY
**Role:** DevOps / SRE
**Salary:** $180k – $220k/yr
**Experience:** 8+ years
**Skills:** Python, Linux, APIs, Infrastructure Automation, Bare-Metal Infrastructure, Gpu Infrastructure, Hpc, Containerization, Orchestration, Pxe/Ipxe, Redfish, Ipmi, Kubernetes, Observability, Sonic
**Posted:** 2026-08-31

> Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.

## Job Description

## Responsibilities

### Infrastructure Software & Automation
- Design, build, and operate production services, APIs, tooling, and automation for large-scale bare-metal and GPU infrastructure.
- Build software systems for infrastructure provisioning, configuration, monitoring, and lifecycle management.
- Develop reliable automation that reduces manual operational work and improves consistency and scalability.
- Build tools and interfaces for programmatic interaction with physical infrastructure.

### Provisioning & Lifecycle Management
- Develop systems for server discovery, provisioning, configuration, validation, and lifecycle management.
- Automate infrastructure bring-up and capacity deployment across large fleets of compute systems.
- Integrate software and automation with hardware management and provisioning systems.
- Improve tooling and workflows across the infrastructure production lifecycle.

### Reliability & Observability
- Build telemetry, logging, and observability capabilities for infrastructure health visibility.
- Develop tools and automation to identify, diagnose, and respond to infrastructure and hardware issues.
- Convert recurring operational issues and hardware failure modes into improvements to software, tooling, and automation.
- Improve the reliability and scalability of infrastructure operations.

### Architecture & Collaboration
- Partner with Network, Infrastructure Operations, Data Center, and Platform Engineering teams to define requirements and deliver infrastructure capabilities.
- Participate in architectural discussions and help define the technical direction of infrastructure software.
- Make pragmatic design decisions balancing reliability, scalability, simplicity, and execution speed.
- Write design documents and technical documentation and contribute to engineering best practices.

## Requirements
- 8+ years of professional software engineering, infrastructure engineering, or related experience.
- Strong software engineering fundamentals and experience building production backend systems in Python or similar object-oriented languages.
- Strong experience with Linux in production environments.
- Experience building APIs, tooling, or automation for managing infrastructure at scale.
- Familiarity with containerization and orchestration concepts.
- Understanding of HPC and bare-metal infrastructure fundamentals, including provisioning and out-of-band management.
- Ability to navigate ambiguity and make pragmatic architecture decisions for long-term reliability and scale.
- Experience in a startup or other fast-paced environment with substantial ownership and autonomy.

## Nice-to-Haves
- Experience with bare-metal hardware troubleshooting and provisioning, including PXE/iPXE, BMC, Redfish, or IPMI, particularly with Dell hardware.
- Experience with GPU servers in bare-metal or virtualized environments.
- Experience with network switches, routers, and firewalls, particularly SONiC switches, Palo Alto firewalls, or Juniper Networks.
- Experience with high-performance storage systems, particularly VAST.
- Experience supporting AI/ML or HPC infrastructure at scale.

## Compensation & Benefits
- Annual base salary: **$180,000–$220,000 USD**.
- Discretionary bonus and meaningful equity component.
- Comprehensive medical, dental, and vision coverage.
- 401(k) matching in the U.S. and pension contributions in the U.K.
- Unlimited PTO, company holidays, and floating holidays.
- Two-week company-wide winter break.
- Paid parental and family leave.
- Annual learning and development allowance.
- Wellness and work-from-home stipends.
- Four-week paid sabbatical after four years of service.
- Flexible schedules and hybrid work model.
- Complimentary in-office meals.

## Similar jobs

- [Senior HPC Storage Engineer](https://hotfix.jobs/jobs/72f52714-20f3-4c62-b022-2468417a44d6) - Runpod - Remote - $180k – $260k/yr
- [Senior Site Reliability Engineer - Linux Systems & Application Observability](https://hotfix.jobs/jobs/f3e1c6e1-7d10-4e98-9005-e0cd58837354) - tastytrade - Chicago, IL - $180k – $200k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/396237d0-dce2-433a-ba7b-1852c20a74eb) - Sprig - San Francisco, CA - $180k – $260k/yr
- [Senior Platform Software Engineer](https://hotfix.jobs/jobs/c1a5fdfa-4570-43d0-9a7f-d777288315e3) - Camber - New York, NY - $180k – $230k/yr
- [Senior Site Reliability Engineer, Colorado Springs](https://hotfix.jobs/jobs/c3d2b727-7792-4fcd-bb4c-3c6d164c72e6) - Onebrief - Colorado Springs, CO - $180k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/a690500a-c87a-4f20-850d-dc77900b33c3
**Canonical:** https://hotfix.jobs/jobs/a690500a-c87a-4f20-850d-dc77900b33c3