Tokens-as-a-Service (Taas) Software Engineer
Builds systems and tooling to measure, monitor, and optimize token throughput from GPU infrastructure for OpenAI workloads. Integrates partner compute environments, benchmarks performance, analyzes tokenomics, and develops operational metrics and dashboards. Requires strong distributed systems and infrastructure engineering experience.
About the job
Key Responsibilities
- Develop systems and tooling to measure, monitor, and improve token throughput across first-party and partner-owned compute environments.
- Support performance benchmarking, tokenomics analysis, and model porting across heterogeneous infrastructure environments.
- Build tooling to integrate external or partner infrastructure into OpenAI's internal compute, observability, and workload management systems.
- Develop and monitor operational metrics including billing, usage, SLAs, utilization, reliability, and throughput.
- Identify bottlenecks across hardware, networking, software, and workload enablement that prevent capacity from becoming productive tokens.
- Partner with compute, infrastructure, networking, finance, and operations teams to translate raw capacity into usable workload-serving capacity.
- Build dashboards, automation, and reporting systems that provide clear visibility into TaaS capacity, performance, and business outcomes.
Qualifications
- Strong software engineering background with experience building systems, tooling, automation, or infrastructure platforms.
- Experience working across compute infrastructure, distributed systems, performance engineering, or production operations.
- Ability to reason about token throughput, utilization, benchmarking, infrastructure efficiency, and workload performance.
- Comfortable integrating external systems or partner environments into internal infrastructure stacks.
- Strong analytical and debugging skills across hardware, networking, software, and operational domains.
Preferred Skills
- Experience with GPU clusters, AI infrastructure, performance benchmarking, or workload optimization.
- Familiarity with model porting, inference/training workloads, token economics, or compute efficiency analysis.
- Experience building monitoring systems for billing, usage, SLAs, utilization, or infrastructure reliability.
- Background in systems engineering, infrastructure software, observability, distributed systems, or platform engineering.
Skills
Distributed Systems, Gpu Clusters, Performance Benchmarking, Observability, Kubernetes, AI Infrastructure, Monitoring Systems, Dashboards, Automation, Infrastructure Integration
Similar jobs
DevOps / SRE jobsBuild and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.
Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.
Build and operate an AI-first CI/CD and agent-operations platform for Salesforce and custom GTM applications. The role focuses on governed releases, approval workflows, observability, rollback, sandboxing, and SOX-compliant auditability.
Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.