Software Engineer, Infrastructure Reliability
Build and operate scalable backend infrastructure supporting high-traffic ChatGPT experiences. The role requires at least four years of software engineering experience, strong production coding skills, and expertise in distributed systems, reliability, performance, and safe deployment practices.
About the job
Responsibilities
- Design, build, and maintain backend systems supporting high-traffic ChatGPT experiences.
- Develop shared services, APIs, and infrastructure that help product teams build and launch new capabilities safely.
- Improve the performance, scalability, and efficiency of production systems as usage and product complexity grow.
- Build and improve systems for asynchronous processing and other large-scale backend workloads.
- Lead architectural improvements and infrastructure migrations while maintaining correctness, compatibility, and safe rollout and rollback.
- Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact.
- Participate in on-call, incident response, and root-cause analysis, turning operational learnings into lasting engineering improvements.
- Collaborate across product and infrastructure teams to ensure new systems are reliable, secure, and production-ready.
- Build automation and tooling that reduce repetitive operational work and improve engineering effectiveness.
Requirements
- 4+ years of professional software engineering experience, including significant experience building backend or distributed systems.
- Strong proficiency in at least one general-purpose programming language and experience writing reliable, maintainable production code.
- Experience designing, building, or improving services, platforms, or shared infrastructure at scale.
- Familiarity with production reliability practices, including monitoring, incident response, root-cause analysis, and operational readiness.
- Experience diagnosing performance, scalability, or reliability issues in production environments.
- Understanding of distributed systems, data storage, concurrency, asynchronous processing, or networking.
- Familiarity with modern deployment, observability, and cloud infrastructure practices.
- Ability to lead complex technical work and collaborate effectively across product and infrastructure teams.
Nice to Have
- Experience with cloud infrastructure, containerized environments, or observability tools.
- Background in backend engineering, distributed systems, platform engineering, or reliability engineering.
Skills
Backend Engineering, Distributed Systems, Production Reliability, APIs, Asynchronous Processing, Concurrency, Data Storage, Networking, Monitoring, Incident Response, Root-Cause Analysis, Observability, Cloud Infrastructure, Containerization, Performance Optimization
Similar jobs
Backend Engineering jobsBuild backend infrastructure and customer-facing workflows that enable developers and AI agents to operate reliably in secure cloud environments. The role requires strong Go and production backend experience, distributed-systems expertise, and practical knowledge of cloud infrastructure, networking, and security.
Build foundational backend and data systems that detect, track, and enforce against harm and abuse on AI platforms. The role requires at least five years of software engineering experience, trust and safety expertise, and strong cross-functional collaboration skills.
Design and lead secure, scalable wallet infrastructure spanning custody, key management, signing, authorization, and recovery. The role requires deep expertise in wallet or cryptographic security infrastructure, distributed systems architecture, and modern blockchain account models.
Build and operate production backend features, APIs, integrations, and services for a large-scale DevSecOps platform. The role requires professional backend development experience, strong database and testing fundamentals, independent feature ownership, and clear asynchronous communication.
Build and maintain Ruby on Rails backend systems for GitLab’s vulnerability management workflows, including security dashboards, reports, and ingestion pipelines. The role requires backend application experience, PostgreSQL knowledge, technical problem-solving, and effective collaboration in a distributed team.