Staff Platform Engineer
Build and operate Ashby’s scalable platform, improving reliability, security, deployment workflows, and developer experience. The role requires strong software engineering skills, infrastructure automation experience, operational judgment, and comfort owning projects end-to-end in a distributed environment.
About the job
Responsibilities
- Scale and extend Ashby’s platform as customer and product demands grow.
- Build infrastructure-as-code and improve how engineering teams interact with infrastructure.
- Own platform projects end-to-end, including scaling, security, reliability, and developer experience.
- Optimize the recruiting DSL-to-SQL compiler and build developer optimization tools.
- Create automated security and privacy guardrails for customer data.
- Enable canary deployments, gradual rollouts, and feature flags while reducing downtime.
- Define SLOs with business and engineering stakeholders and implement corresponding SLIs.
- Ensure communication with external services supports retries and circuit breakers.
- Build infrastructure for event-driven architecture and a data warehouse.
- Perform infrastructure updates, security enforcement, database optimization, Kubernetes debugging, and application troubleshooting.
- Contribute to on-call operations in a follow-the-sun model.
- Improve developer tooling and communicate best practices through written documentation and code.
Requirements
- Professional experience building infrastructure at a later-stage company.
- Experience handling large volumes of data and understanding infrastructure’s effect on customer experience.
- Experience automating provisioning, monitoring, and release processes.
- Strong software engineering and coding ability; this role includes code reviews and code changes.
- Comfort making independent technical decisions and delivering projects end-to-end.
- Ability to evaluate reliability and operational risk.
- Comfort with SQL and data modeling.
- Ability to debug Kubernetes issues and investigate application traces.
- Strong written communication skills and comfort working asynchronously.
Nice-to-haves
- Experience with TypeScript, Node.js, React, Apollo GraphQL, Postgres, Redis, Datadog, Sentry, and AWS.
- Experience with infrastructure-as-code, SRE practices, event-driven architectures, data warehouses, canary deployments, gradual rollouts, feature flags, SLOs, and SLIs.
Compensation and benefits
- Annual salary range: $154,000–$250,000.
Skills
TypeScript, Node.js, React, Apollo Graphql, Postgres, Redis, Datadog, Sentry, AWS, Kubernetes, SQL, Infrastructure As Code, SLOs, Slis, Feature Flags
Similar jobs
DevOps / SRE jobsOwn and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.
Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Leads the design, operation, and evolution of Flexport’s cloud infrastructure, platform tooling, observability, and incident response systems. Requires 10+ years of software, SRE, or infrastructure engineering experience, deep AWS expertise, and strong Terraform and automation skills.