Build and operate reliable, scalable infrastructure and automation for OpenAI's research workloads and data systems (acquisition, processing, ingest, search). Requires strong systems and distributed systems experience, Kubernetes, Linux, networking, and software engineering to improve reliability and reduce operational toil.
255k – 405k/yr
On-site5+ YOEData Engineering
About the role
Responsibilities
Build and operate reliable infrastructure for research workloads and research-facing services.
Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services.
Improve cluster bootstrapping, provisioning, automation, and deployment workflows.
Debug issues across networking, compute, storage, orchestration, and service reliability layers.
Build software and automation that reduce manual operational work and improve system reliability.
Partner closely with researchers, infrastructure engineers, and service owners to understand system needs and translate them into durable solutions.
Help evolve existing infrastructure toward more scalable, maintainable, and standard patterns.
Take ownership of critical systems and drive work independently from problem definition through execution.
Requirements
Strong systems fundamentals and understanding of how infrastructure scales in practice.
Comfortable with Linux, networking, Kubernetes, provisioning, and distributed systems operations.
Ability to write software to automate, debug, and improve infrastructure systems.
Strong execution mindset and ability to independently drive ambiguous infrastructure work.
Enjoy supporting a wide surface area of systems, from research tooling to platform services.
Pragmatic about when to build custom systems versus using existing, well-supported tools.
Care about building reliable systems that make researchers faster and reduce operational friction.
Nice-to-Haves
Experience with PXE boot, cluster provisioning, bare-metal infrastructure, or large-scale fleet management.
Experience operating Kubernetes or similar orchestration systems at scale.
Experience with infrastructure-as-code, CI/CD, observability, or deployment automation.
Experience supporting search infrastructure, data platforms, ingest systems, or large-scale research workflows.
Experience with Git-based workflows and internal developer tooling.
Prior experience in environments where reliability, scale, and speed all matter.
Skills
KubernetesLinuxNetworkingDistributed SystemsAutomationInfrastructure As CodeCI/CDObservabilityPythonDebugging
Designs and implements dataset infrastructure for OpenAI's large-scale LLM training stack, including standardized APIs for multimodal data, scaling pipelines across GPU fleets, and performance debugging. Requires strong distributed systems experience and collaboration with researchers.
250k – 380k/yr
On-siteData Engineering
Dev Ops, Facilities Pipeline
FluidstackAustin, TX +3
Build and operate high-frequency facilities telemetry data pipelines (power, cooling, BMS, sensors) for scaling AI data centers. Stand up ingestion/streaming infrastructure, automate deployments with IaC, and own end-to-end reliability for gigawatt-scale operations.
269k – 317k/yr
On-site5+ YOEData Engineering
Software Engineer, Facilities Pipeline
FluidstackAustin, TX +3
Build and own the facilities data pipeline for AI data center telemetry, including ingestion from industrial protocols (BACnet, Modbus, OPC UA), data quality tooling, and serving clean APIs/datasets for dashboards, controls, and ML. Requires production data pipeline experience with on-call ownership, industrial protocol integration, and full-stack debugging.
269k – 317k/yr
On-site5+ YOEData Engineering
Data Engineer
FluidstackAustin, TX +3
Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.
269k – 317k/yr
On-site5+ YOEData Engineering
Data Operations Manager, Human Data
AnthropicSan Francisco, CA +1
Data Operations Manager responsible for building and scaling data strategies, vendor partnerships, and high-quality data pipelines to advance frontier AI research in RLHF, safety, tool use, and agentic systems at Anthropic. Requires 3+ years operations/PM experience, strong project management, data analysis skills, and passion for AI safety.