Senior Site Reliability Engineer
Nectar Social is seeking a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of their production systems. This role involves defining SLOs, leading incident response, improving infrastructure performance, and partnering with engineering teams to embed reliability into system design.
About the job
The Role
We're looking for a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of the production systems that power Nectar's platform. We run high-volume data ingestion pipelines and real-time AI agents on top of a fast-growing customer base, and we need a seasoned SRE to help us scale these systems safely and keep them running flawlessly.
As one of our first dedicated SREs, you'll have outsized impact and ownership. You'll define how we measure, operate, and harden our infrastructure -- establishing the reliability foundations that the rest of the engineering team builds on as we scale.
What You'll Be Doing
- Own the reliability and scalability of our production systems as they handle rapidly growing volumes of social data and real-time AI workloads
- Define and drive SLOs, SLIs, and error budgets, and build the observability, alerting, and on-call practices to support them
- Lead incident response and blameless postmortems, then turn what we learn into systemic improvements that prevent recurrence
- Improve performance, cost efficiency, and capacity planning across our cloud infrastructure as the platform scales
- Harden our infrastructure-as-code, deployment, and CI/CD pipelines for resilience and repeatability
- Partner with engineering teams to embed reliability into system design and raise the operational bar across the org
What We're Looking For
- 5+ years of experience operating production systems as an SRE, infrastructure, or platform engineer
- Experience scaling databases, data infrastructure, or complex production platforms under significant load
- Hands-on expertise with cloud infrastructure (AWS or similar) and infrastructure-as-code tooling
- Solid programming skills for building automation, tooling, and operational services
- Comfortable operating in fast-moving startup environments with high ownership and autonomy
- A reliability-first mindset balanced with pragmatism about velocity and cost
Bonus Points
- Experience standing up or maturing an SRE practice at an early-stage or rapidly scaling company
- Familiarity with our tech stack: AWS, Pulumi, Postgres, ClickHouse, Turbopuffer, or Temporal
- Background in capacity planning, performance engineering, or cost optimization at scale
What We Offer
- Competitive compensation and early equity
- Health, vision, and dental benefits + 401(k) match
- Clear career growth opportunities as the company scales
- Free lunch in the heart of University Ave. in Palo Alto
- Deep exposure to cutting-edge AI tooling and the opportunity to shape how brands use it
- A collaborative, ambitious team defining a new category of AI-native marketing infrastructure
Skills
AWS, Pulumi, Postgres, ClickHouse, Temporal, CI/CD, Infrastructure-As-Code, System Design, Capacity Planning, Performance Engineering
Similar jobs
DevOps / SRE jobsBuild and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.