Datacenter Hardware Operations Technician, AI Compute Infrastructure - Stargate
Senior on-site hardware operations lead responsible for server, GPU, storage, and rack infrastructure reliability at OpenAI's AI campuses. Requires 8+ years of large-scale datacenter hardware experience and strong troubleshooting and cross-functional leadership skills.
About the job
Key Responsibilities
- Serve as OpenAI’s senior on-site hardware operations lead for server, GPU, storage, and rack-level infrastructure
- Drive technical triage and resolution of complex hardware failures impacting production systems
- Partner with Fleet Health Engineering to investigate recurring hardware issues, identify failure patterns, and improve fleet reliability
- Lead root cause analysis (RCA) efforts for critical hardware incidents and develop corrective and preventive action plans
- Collaborate with Oracle operations teams and OEM vendors to coordinate repairs, replacements, upgrades, and hardware lifecycle activities
- Establish and continuously improve hardware maintenance procedures, operational runbooks, and troubleshooting standards
- Analyze hardware failure trends and operational metrics to identify reliability risks and improvement opportunities
- Support new hardware introductions, validation activities, and production readiness reviews
- Coordinate spare parts strategy and inventory planning with supply chain and operations teams
- Partner with Hardware Engineering, Manufacturing, and Infrastructure teams to provide field feedback that improves future platform designs
- Develop scalable operational standards and best practices that can be deployed across future Stargate campuses
- Mentor technicians and partner teams on advanced troubleshooting methodologies and hardware operational excellence
Qualifications
- 8+ years of experience supporting large-scale datacenter hardware infrastructure, with experience in a senior technician, sustaining engineering, or hardware operations leadership role
- Deep expertise with server platforms, GPU systems, storage infrastructure, rack integration, and datacenter hardware architecture
- Strong experience diagnosing complex hardware failures and leading repair efforts in production environments
- Experience conducting root cause analysis and driving long-term corrective actions
- Strong understanding of hardware reliability engineering principles and fleet-health management
- Proven ability to partner effectively across engineering, operations, manufacturing, and vendor organizations
- Comfortable operating independently in high-priority production environments with significant operational responsibility
- Excellent written and verbal communication skills with the ability to influence technical and operational decisions
- Experience developing operational processes, maintenance standards, and technical documentation
- Ability to travel occasionally to support new campus deployments and operational readiness activities
Preferred Qualifications
- Experience supporting large-scale GPU clusters or AI/ML infrastructure environments
- Familiarity with fleet health systems, telemetry platforms, and hardware monitoring tools
- Experience with failure analysis methodologies such as FRACAS, RCCA, 5-Why, Fishbone, or FMEA
- Knowledge of Linux system administration and hardware validation workflows
- Experience supporting hyperscale datacenter operations or HPC environments
- Familiarity with server manufacturing, rack integration, or NPI-to-sustaining transitions
- Industry certifications such as CompTIA Server+, OEM hardware certifications, or equivalent experience
- Experience applying Environmental Health and Safety (EHS) practices in mission-critical datacenter environments
Skills
Server Platforms, Gpu Systems, Storage Infrastructure, Rack Integration, Datacenter Hardware Architecture, Root Cause Analysis, Hardware Reliability Engineering, Fleet Health Management, Linux System Administration, Hardware Validation
Similar jobs
Hardware Engineering jobsOwn and optimize critical facility infrastructure supporting AI data center operations, including electrical, mechanical, cooling, and building management systems. The role requires at least three years of facilities or critical-infrastructure experience, strong troubleshooting skills, and familiarity with applicable safety and engineering standards.
Leads development and validation of steady-state and dynamic process models for integrated mineral refining systems. Requires 8+ years of process simulation experience, engineering expertise in transport and thermodynamics, and proficiency with industrial simulation tools.
Leads supplier qualification, audits, NPI quality, corrective actions, and continuous improvement for manufacturing suppliers. Requires a bachelor’s degree, 3+ years of relevant experience, strong quality-methodology expertise, engineering drawing literacy, and cross-functional influence.
Leads the design, construction, deployment, and live operations of a global ground-station network from the Los Angeles headquarters. The role requires 5+ years of relevant engineering experience, a bachelor's degree, hands-on field expertise, and frequent domestic and international travel.
Builds and leads supplier quality systems for complex RF, PCBA, mechanical, and electromechanical hardware used in satellite ground stations. The role requires 5+ years in supplier quality or development, supplier qualification expertise, and experience with aerospace-quality standards and manufacturing ramp-up.