# Software Engineer, Infrastructure Reliability

**Company:** [OpenAI](https://hotfix.jobs/companies/openai)
**Location:** London, United Kingdom
**Role:** Backend Engineering
**Salary:** $255k – $405k/yr
**Experience:** 4+ years
**Skills:** Backend Engineering, Distributed Systems, Production Reliability, APIs, Asynchronous Processing, Concurrency, Data Storage, Networking, Monitoring, Incident Response, Root-Cause Analysis, Observability, Cloud Infrastructure, Containerization, Performance Optimization
**Posted:** 2026-07-28

> Build and operate scalable backend infrastructure supporting high-traffic ChatGPT experiences. The role requires at least four years of software engineering experience, strong production coding skills, and expertise in distributed systems, reliability, performance, and safe deployment practices.

## Job Description

## Responsibilities
- Design, build, and maintain backend systems supporting high-traffic ChatGPT experiences.
- Develop shared services, APIs, and infrastructure that help product teams build and launch new capabilities safely.
- Improve the performance, scalability, and efficiency of production systems as usage and product complexity grow.
- Build and improve systems for asynchronous processing and other large-scale backend workloads.
- Lead architectural improvements and infrastructure migrations while maintaining correctness, compatibility, and safe rollout and rollback.
- Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact.
- Participate in on-call, incident response, and root-cause analysis, turning operational learnings into lasting engineering improvements.
- Collaborate across product and infrastructure teams to ensure new systems are reliable, secure, and production-ready.
- Build automation and tooling that reduce repetitive operational work and improve engineering effectiveness.

## Requirements
- 4+ years of professional software engineering experience, including significant experience building backend or distributed systems.
- Strong proficiency in at least one general-purpose programming language and experience writing reliable, maintainable production code.
- Experience designing, building, or improving services, platforms, or shared infrastructure at scale.
- Familiarity with production reliability practices, including monitoring, incident response, root-cause analysis, and operational readiness.
- Experience diagnosing performance, scalability, or reliability issues in production environments.
- Understanding of distributed systems, data storage, concurrency, asynchronous processing, or networking.
- Familiarity with modern deployment, observability, and cloud infrastructure practices.
- Ability to lead complex technical work and collaborate effectively across product and infrastructure teams.

## Nice to Have
- Experience with cloud infrastructure, containerized environments, or observability tools.
- Background in backend engineering, distributed systems, platform engineering, or reliability engineering.

## Similar jobs

- [Product Engineer, Ona](https://hotfix.jobs/jobs/fced2e7b-d946-44bb-8c1c-50640be580c3) - OpenAI - San Francisco, CA - $255k – $445k/yr
- [Software Engineer, Integrity Foundations](https://hotfix.jobs/jobs/3599a04a-74a1-4881-b7dd-ca086d3c67c3) - OpenAI - London, United Kingdom - $221k – $370k/yr
- [Software Engineer - Wallets](https://hotfix.jobs/jobs/32955ac3-4c1c-4a94-83a2-7e6f18965180) - Rain - Remote - $220k – $270k/yr
- [Intermediate Backend Engineer](https://hotfix.jobs/jobs/8113217a-26e2-4d4b-86f1-55fa954ef043) - GitLab - Remote - $115k – $194k/yr
- [Intermediate Software Engineer, Security Factory: Vulnerability Management](https://hotfix.jobs/jobs/9f09cc9e-0cdb-435d-8308-8947f044ea04) - GitLab - Remote - $115k – $173k/yr

**Apply:** https://hotfix.jobs/jobs/015adee7-58ba-476f-a5e0-0f62c1efdaf2
**Canonical:** https://hotfix.jobs/jobs/015adee7-58ba-476f-a5e0-0f62c1efdaf2