Software Engineer, ChatGPT Infrastructure
Design and build infrastructure platforms for ChatGPT, focusing on scalability, reliability, and developer productivity. Requires experience with large-scale distributed systems, performance optimization, and creating reusable abstractions for engineering teams.
About the job
Where You Can Have Impact
- Platform foundations & frameworks: core libraries, service frameworks, and shared components.
- Scalability & performance primitives: reduce tail latency, improve throughput, keep costs predictable.
- Reliability guardrails: rate limiting, load shedding, dependency isolation, backpressure, safe fallbacks.
- Developer productivity via golden paths: paved roads for common workflows.
- Observability & debugging systems: instrumentation, metrics, investigative tooling.
- Safe change management: deployment systems, progressive delivery, fast rollback.
- Interface and contract design: clean APIs and stable contracts.
What You’ll Do
- Build and evolve infrastructure platforms that many engineers and services depend on.
- Translate messy real-world constraints into clean abstractions: simple APIs, enforceable contracts, safe defaults.
- Drive improvements in reliability and performance through principled design, measurement, and iterative hardening.
- Partner across engineering and product to identify systemic pain points and turn them into reusable solutions.
- Own outcomes end-to-end: design → implementation → rollout → operational maturity.
Qualifications
Minimum Qualifications
- Experience building and operating large-scale distributed systems in production (high throughput, concurrency, and failure handling).
- Strong fundamentals in systems design, including caching, consistency, queueing/backpressure, and resilient dependency management.
- Ability to reason about performance (latency distributions, tail behavior, bottlenecks) and translate that into concrete engineering work.
- Track record of building platforms or shared infrastructure that improves velocity and correctness for other teams.
- Strong communication and collaboration skills—aligning on interfaces, navigating tradeoffs, and driving cross-team execution.
Preferred Qualifications
- Experience designing paved roads / golden paths (frameworks, libraries, self-serve tooling) that shape engineering behavior at scale.
- Deep understanding of reliability techniques: graceful degradation, circuit breakers, load shedding, rate limiting, and fault isolation.
- Experience building systems for safe iteration: progressive delivery, correctness checks, automated rollout gates, and production validation.
- Strong instincts for API and contract design—how to create interfaces that are stable, evolvable, and hard to misuse.
- Prior work that demonstrates “force multiplier” impact: enabling many teams through a small number of well-chosen primitives.
Skills
Distributed Systems, System Design, Caching, Backpressure, Rate Limiting, Load Shedding, Observability, API Design, Progressive Delivery, Tail Latency Optimization
Similar jobs
Backend Engineering jobsBuild backend infrastructure and customer-facing workflows that enable developers and AI agents to operate reliably in secure cloud environments. The role requires strong Go and production backend experience, distributed-systems expertise, and practical knowledge of cloud infrastructure, networking, and security.
Build and operate Host Assurance services and host software that establish trust in bare-metal and virtual machine infrastructure through secure bootstrap, identity, attestation, and verification. The role requires production software engineering experience across reliable systems, platform or infrastructure security, and host-system boundaries.
Build and own Baseten’s identity and authorization platform, including fine-grained permissions, credential systems, and enterprise administration. The role requires backend systems experience, production authorization expertise, and the ability to operate secure multi-tenant systems at scale.
Build and operate scalable backend services, APIs, and real-time data systems while contributing to architecture, reliability, and product-driven solutions. Requires a relevant master’s degree with 3 years of experience or a bachelor’s degree with 5 years, plus broad distributed systems and programming expertise.
Build and operate backend and financial infrastructure supporting pricing, payments, billing, subscriptions, entitlements, and invoicing. The role requires 5+ years of software engineering experience with distributed systems, transactional workflows, and highly reliable, auditable platforms.