Member of Technical Staff
Site Reliability Engineer building and operating scalable, reliable infrastructure for an AI-powered accounting platform at a fast-growing startup. Requires 5+ years production infrastructure experience, strong software engineering skills, IaC, observability, and on-call ownership in an onsite NYC role.
About the job
What You’ll Do
- Architect, build, and operate reliable, scalable, and secure infrastructure for our production systems.
- Own cloud infrastructure across compute, storage, and networking, optimizing for availability, performance, and cost efficiency.
- Design and maintain CI/CD pipelines, infrastructure-as-code, and automation to improve developer velocity and system reliability.
- Lead incident response efforts, including on-call rotations, incident coordination, postmortems, and root cause analyses.
- Partner closely with product and engineering teams to balance reliability, performance, cost, and speed of development.
- Automate operational workflows such as capacity planning, safe rollouts, graceful degradation, and data access controls.
- Provide technical leadership and mentorship, helping shape the culture and standards of the infrastructure team.
What You’ll Bring
- 5+ years of experience building and operating production infrastructure at scale.
- Strong software engineering fundamentals and proficiency in at least one programming language.
- Deep understanding of cloud infrastructure, networking, databases, and security principles.
- Experience with CI/CD systems, containerization, and modern infrastructure automation.
- Comfort operating in ambiguous environments and taking ownership end-to-end.
- Experience with similar stack (or ability to learn unfamiliar technologies): Infrastructure-as-Code tools such as Terraform or CloudFormation; Observability and incident management tools (e.g., OpenTelemetry, Prometheus/Grafana, BetterStack, PagerDuty, SLOs/error budgets); Data and systems powering analytics and customer-facing workloads, especially in serverless environments (e.g., Neon, Modal).
Nice-to-Haves
- Excited about working with coding agents.
- First principles and systems thinking.
- Thinking beyond the engineering, understanding domain deeply.
- Ownership of outcomes and agency.
Benefits
- Health & Wellness: Premium Medical, Dental, and Vision coverage; Life Insurance; and 6 coaching & 6 therapy sessions through Spring Health.
- Time off: Unlimited PTO + 12 paid company holidays.
- In-Office Perks: Daily meal stipends, a fully stocked kitchen, and $300 toward your custom desk setup.
- Financial Benefits: Pre-tax commuter benefits and 401(k) retirement plan.
- Team Culture: Monthly office activities and frequent optional team happy hours.
- Parental Leave
Skills
Terraform, CloudFormation, OpenTelemetry, Prometheus, Grafana, Pagerduty, SLOs, CI/CD, Infrastructure As Code, Containerization, Cloud Infrastructure, Networking, Databases, Security
Similar jobs
DevOps / SRE jobsThe DevOps Engineer will design and operate AWS and hybrid infrastructure, improve CI/CD reliability, and strengthen disaster recovery and business continuity. The role requires 5+ years of cloud infrastructure experience, strong AWS proficiency, and hands-on infrastructure-as-code expertise.
Operates and evolves foundational networking, compute, Kubernetes, and ingress infrastructure for PagerDuty’s real-time platform. Requires 3+ years in SRE, DevOps, or platform engineering, with Linux production operations, cloud infrastructure, programming, and Infrastructure as Code experience.
Owns secure, scalable Azure infrastructure for healthcare applications, including cloud migrations, Terraform-based automation, CI/CD pipelines, monitoring, and compliance. Requires 3–5+ years of Azure experience and strong DevOps and cloud-security expertise.
Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.
Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.