Staff Platform Engineer, Americas
Staff Platform Engineer builds and scales infrastructure, optimizes compilers and databases, implements deployment tools like canary deploys and feature flags, and ensures reliability with SLOs/SLIs on AWS/Kubernetes. Requires strong coding skills in TypeScript/Node.js and handling diverse infra challenges end-to-end.
About the job
Responsibilities
- Optimize homegrown ultra-dynamic recruiting DSL-to-SQL compiler and create tools to help developers.
- Create automated guardrails for security and privacy of customer data.
- Help developers ship features fast through canary deploys, gradual rollouts, and feature flags while managing complexity and reducing downtime.
- Work with business and engineering to define SLOs and implement SLIs.
- Ensure communication with external services supports retries and circuit-breakers.
- Implement infrastructure for event-driven architecture and data warehouse.
Requirements
- Build systems that are mature, reliable, flexible, and approachable.
- Comfortable evaluating risk and comfortable coding (reviewing and submitting code changes).
- Experience building infrastructure at scale with millions of data points, automating provisioning, monitoring, and release processes.
- Handle diverse problems: infrastructure updates, security, database optimization, Kubernetes debugging, Typescript traces.
- Strong in SQL; dive into reports and advise on performant data models.
- Async written communication for sharing tooling and best practices.
- Own projects end-to-end with accountability.
Technology Stack
- TypeScript (frontend & backend)
- Node.js
- React
- Apollo GraphQL
- Postgres
- Redis
- Datadog
- Sentry
- AWS
- Kubernetes
Compensation
$190,000 - $275,000 USD
Skills
TypeScript, Node.js, React, Apollo Graphql, Postgres, Redis, Kubernetes, AWS, Datadog, Sentry, SQL
Similar jobs
DevOps / SRE jobsLeads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.
Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.