Own hands-on Layer 1 activation, troubleshooting, and automation for WAN, fiber, carrier, and cloud interconnect circuits in OpenAI's GPU fleet data centers. Requires 3+ years network operations experience, optical fault isolation skills, provider coordination, and practical automation with scripts/APIs/LLM tools.
157k – 221k/yr
On-site3+ YOEDevOps / SRE
About the role
Responsibilities
Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence.
Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports.
Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching, incompatible optics or media, wavelength/power-budget problems, breakout, speed, and FEC/PCS mismatches.
Guide remote hands, carrier technicians, CSP operations, and colocation teams through targeted inspection/cleaning, VFL, optical power, OTDR/OLTS, approved loopback, and per-lane DOM/DDM tests.
Verify service acceptance on both sides: correct circuit and port, admin/operational state, compatible optics, light levels, speed/FEC, clean counters, stable link, and Layer 2/3 handoff readiness.
Own tickets and troubleshooting bridges through resolution; identify the blocker and A-side/Z-side owner, drive follow-ups, escalate with evidence, and prevent premature closure.
Execute smart-hands and change work safely: name the exact circuit, port, and fiber pair; protect adjacent live services; get approval before intrusive tests, repatches, shutdowns, or loopbacks; and maintain rollback.
Deliver complete production handoffs with verified mappings, provider/CSP references, test results, interface evidence, incident history, risks, and clear ownership; improve runbooks and prevent repeat faults.
Automate the full circuit-bring-up lifecycle: structure circuit, port, and patch-panel data; ingest telemetry and provider tickets; generate LLM-assisted diagnostics and technician-ready instructions; track ownership and evidence; and preserve explicit human approval for intrusive work.
Qualifications
3+ years of hands-on network operations, carrier activation, data-center network deployment, field engineering, optical transport, or related infrastructure experience; equivalent practical experience is valued.
Can trace an end-to-end circuit across cross-connects, fiber pairs, patch panels, carrier demarcations, optics, and both A-side and Z-side device ports.
Troubleshoot optical faults using inspection/cleaning tools, optical power meters, VFL, OLTS/OTDR evidence, DOM/DDM, interface status, and provider test results.
Understand single-mode/multimode fiber, LC and MPO/MTP connectors, duplex polarity, TX/RX paths, wavelength and power budgets, and 100G/400G/800G optics.
Have worked with carriers, colocation providers, remote hands, AWS Direct Connect, Azure ExpressRoute, or Google Cloud Interconnect, and persist through multi-party ownership gaps.
Can build practical automation with scripts, APIs, and structured data, and safely translate LLM-generated analysis into validated actions for data-center technicians.
Operate and maintain large-scale Ethernet fabrics for OpenAI's AI GPU clusters and infrastructure. Troubleshoot incidents, execute changes, perform RCAs, build observability, and automate operations across data centers while partnering with engineering and vendor teams.
157k – 221k/yrOn-site5+ YOEDevOps / SRE
Software Engineer, DevInfra
MixpanelSan Francisco, CA
DevInfra engineer building and maintaining developer tooling, Kubernetes infrastructure, CI/CD pipelines, and AI-powered automation on GCP to accelerate engineering velocity.
158k – 213k/yrRemote3+ YOEDevOps / SRE
Database Administrator (DBA)
The Voleon GroupBerkeley, CA +1
Manages production and development relational databases (primarily PostgreSQL), handling operations like backups, performance tuning, monitoring, and automation. Collaborates with infrastructure teams on storage, OS, and networking for optimal reliability and performance.
155k – 195k/yrRemoteDevOps / SRE
DevOps Engineer (USA or Canada)
PanoptoUnited States
Mid-level DevOps Engineer modernizes legacy CI/CD pipelines, implements IaC with Terraform/CloudFormation, and drives automation using Docker, Kubernetes, and scripting in Python/C#/Bash. Requires 3-5 years experience in DevOps/SRE and a bachelor's degree.
155k – 175k/yrRemote3+ YOEDevOps / SRE
🛠️ Platform Engineer
PartifulNew York, NY
Builds and maintains scalable infrastructure on GCP, including serverless systems, CI/CD pipelines, and observability tools for a high-growth social events app. Requires 3+ years in infrastructure/backend engineering with serverless experience and strong ownership mindset.