Skip to content
OpenAIOpenAISan Francisco, CA

Software Engineer

Build and operate reliable, scalable infrastructure and automation for OpenAI's research workloads and data systems (acquisition, processing, ingest, search). Requires strong systems and distributed systems experience, Kubernetes, Linux, networking, and software engineering to improve reliability and reduce operational toil.

255k – 405k/yr
On-site5+ YOEData Engineering

About the role

Responsibilities

  • Build and operate reliable infrastructure for research workloads and research-facing services.
  • Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services.
  • Improve cluster bootstrapping, provisioning, automation, and deployment workflows.
  • Debug issues across networking, compute, storage, orchestration, and service reliability layers.
  • Build software and automation that reduce manual operational work and improve system reliability.
  • Partner closely with researchers, infrastructure engineers, and service owners to understand system needs and translate them into durable solutions.
  • Help evolve existing infrastructure toward more scalable, maintainable, and standard patterns.
  • Take ownership of critical systems and drive work independently from problem definition through execution.

Requirements

  • Strong systems fundamentals and understanding of how infrastructure scales in practice.
  • Comfortable with Linux, networking, Kubernetes, provisioning, and distributed systems operations.
  • Ability to write software to automate, debug, and improve infrastructure systems.
  • Strong execution mindset and ability to independently drive ambiguous infrastructure work.
  • Enjoy supporting a wide surface area of systems, from research tooling to platform services.
  • Pragmatic about when to build custom systems versus using existing, well-supported tools.
  • Care about building reliable systems that make researchers faster and reduce operational friction.

Nice-to-Haves

  • Experience with PXE boot, cluster provisioning, bare-metal infrastructure, or large-scale fleet management.
  • Experience operating Kubernetes or similar orchestration systems at scale.
  • Experience with infrastructure-as-code, CI/CD, observability, or deployment automation.
  • Experience supporting search infrastructure, data platforms, ingest systems, or large-scale research workflows.
  • Experience with Git-based workflows and internal developer tooling.
  • Prior experience in environments where reliability, scale, and speed all matter.

Skills

KubernetesLinuxNetworkingDistributed SystemsAutomationInfrastructure As CodeCI/CDObservabilityPythonDebugging
OpenAI

Software Engineer, Data Infrastructure - Research

OpenAISan Francisco, CA

Designs and implements dataset infrastructure for OpenAI's large-scale LLM training stack, including standardized APIs for multimodal data, scaling pipelines across GPU fleets, and performance debugging. Requires strong distributed systems experience and collaboration with researchers.

250k – 380k/yr
On-siteData Engineering
Fluidstack

Dev Ops, Facilities Pipeline

FluidstackAustin, TX +3

Build and operate high-frequency facilities telemetry data pipelines (power, cooling, BMS, sensors) for scaling AI data centers. Stand up ingestion/streaming infrastructure, automate deployments with IaC, and own end-to-end reliability for gigawatt-scale operations.

269k – 317k/yr
On-site5+ YOEData Engineering
Fluidstack

Software Engineer, Facilities Pipeline

FluidstackAustin, TX +3

Build and own the facilities data pipeline for AI data center telemetry, including ingestion from industrial protocols (BACnet, Modbus, OPC UA), data quality tooling, and serving clean APIs/datasets for dashboards, controls, and ML. Requires production data pipeline experience with on-call ownership, industrial protocol integration, and full-stack debugging.

269k – 317k/yr
On-site5+ YOEData Engineering
Fluidstack

Data Engineer

FluidstackAustin, TX +3

Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.

269k – 317k/yr
On-site5+ YOEData Engineering
Anthropic

Data Operations Manager, Human Data

AnthropicSan Francisco, CA +1

Data Operations Manager responsible for building and scaling data strategies, vendor partnerships, and high-quality data pipelines to advance frontier AI research in RLHF, safety, tool use, and agentic systems at Anthropic. Requires 3+ years operations/PM experience, strong project management, data analysis skills, and passion for AI safety.

270k – 365k/yr
Hybrid3+ YOEData Engineering