ChatGPT Performance Engineer
Performance Engineer optimizes infrastructure and application performance for ChatGPT and OpenAI API, focusing on latency, throughput, and efficiency at scale. Requires 7+ years in high-scale systems with expertise in profiling, tracing, and cross-layer optimizations.
About the job
Responsibilities
- Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond.
- Develop tooling and metrics that provide deep observability into system performance.
- Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements.
- Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale.
- Lead investigations into high-impact performance regressions or scalability issues in production.
- Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems.
Requirements
- 7+ years of experience in software engineering with a strong track record in performance or reliability of high-scale distributed systems.
- Deeply comfortable with performance profiling tools and tracing systems.
- Experience optimizing performance across one or more layers of the stack (e.g., database, networking, storage, application runtime, GC tuning, Python/Golang internals, GPU utilization).
- Strong understanding of OS internals, scheduling, memory management, and IO patterns.
- Contributed to observability, benchmarking, or performance-focused infrastructure at scale.
- Demonstrated success navigating ambiguity and aligning stakeholders around performance goals.
- Value simplicity, rigor, and collaboration when solving complex systems problems.
Skills
Python, Go, Performance Profiling, Distributed Systems, Gpu Utilization, Observability, Tracing Systems, Networking, Storage Optimization, Os Internals
Similar jobs
DevOps / SRE jobsLeads operations outcomes for partner-operated data center sites, directing vendors, defining operational standards, and ensuring deployment velocity, availability, repair performance, and incident response. Requires 8+ years in data center or infrastructure operations, vendor oversight experience, and hands-on server, network, and rack-level expertise.
Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.