Staff+ Software Engineer, Platform
Staff-level software engineer builds and scales platform infrastructure across teams, including dev tools, service infra, multicloud, auth, connectivity, API distributability, and ML adaptation systems. Requires 8+ years full-stack experience with Staff leadership, focusing on robust, scalable solutions in fast-paced AI environment.
About the job
What you'll do
Platform Acceleration
- Architect and optimize critical development infrastructure including dev environments, observability, and CI/CD pipelines
- Partner with product teams to understand workflows and eliminate friction points
Service Infra
- Build and maintain core infrastructure: service mesh, observability systems, deployment pipelines, shared libraries
- Enable product teams to build and operate reliable services at scale
Multicloud
- Build infrastructure for multi-cloud providers: cloud-agnostic tooling, cross-cloud networking, multi-region deployments
Auth & Identity
- Build scalable solutions for user authentication, authorization, RBAC, SSO
- Work with product teams, security, support, trust & safety
Connectivity
- Own MCP proxy, OAuth/token management, MCP spec, Python/TypeScript SDKs
- Handle token refresh at scale, admin controls, proxy infrastructure
API Distributability
- Transform Claude API into cloud-native managed product: cross-cloud, on-prem, enterprise security/compliance
Platform Intelligence
- Build training systems for customer-specific Claude adaptation
- Work on ML training infra, production ML pipelines
You might be a good fit if you
- Have 8-10+ years of practical full-stack engineering experience, ideally 2+ years at Staff level
- Led design/delivery of complex user-facing products across full stack
- Technical expert in modern frontend/backend (e.g., React, TypeScript)
- Product-focused: robust, scalable, easy-to-use solutions
- Experience in fast-moving environments, building 0-to-1 products
- Invest in peer mentorship/growth
- Drive cross-team alignment, influence without authority
- Established engineering standards, component architectures, best practices
- Thrive in fast-paced, ambiguous environments
Strong candidates may also
- Technical lead/architect for foundational platform systems
- Designed/scaled billing/payments at high volumes
- Containerization, secure execution environments
- Identity/access management (auth, SSO, RBAC) at enterprise scale
- ML/AI systems, LLM inference, model serving
- Multi-cloud, cross-region architectures
- API design focused on developer experience
Skills
React, TypeScript, CI/CD, Kubernetes, OAuth, Service Mesh, Observability, Multi-Cloud, RBAC, SSO, Ml Training, Llm Inference, Api Gateways, Python, Typescript Sdk
Similar jobs
DevOps / SRE jobsBuild and operate portable infrastructure that enables Claude to run reliably across multiple cloud providers and accelerator platforms. The role requires 8+ years of distributed-systems experience, multi-cloud architecture expertise, production programming, Kubernetes, and Infrastructure as Code proficiency.
Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.