Site Reliability Engineer
Build the Site Reliability Engineering function from the ground up at Forward, defining SLOs, building observability infrastructure, leading incident response, and embedding reliability into the SDLC for their complex SaaS platform. Requires 6+ years SRE/DevOps experience, strong networking and Kubernetes skills, and a track record maturing SRE practices.
About the job
What You'll Own
- Define and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use
- Drive the reliability and operational excellence of the Forward SaaS platform
- Build and maintain observability infrastructure — logging, metrics, tracing, and alerting — so the team always knows what's happening before customers do
- Lead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twice
- Partner with engineering teams to embed reliability thinking into the SDLC — capacity planning, load testing, chaos engineering, and production readiness reviews
- Help define and build the SRE team as the company scales — this is a foundational hire with a path to leadership
What We're Looking For
- 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment
- Proven experience building or significantly maturing an SRE function — not just operating within one someone else built
- Strong fundamentals in networking — TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus
- Hands-on experience with Kubernetes and container orchestration in production environments
- Deep proficiency with observability tooling — Prometheus, Grafana, Datadog, Splunk, or similar
- Strong scripting and automation skills in Python, Bash, or similar
- Experience with cloud platforms — AWS, GCP, or Azure — including infrastructure as code (Terraform, Ansible, or equivalent)
- Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements
- Ability to communicate clearly with both engineering teams and non-technical stakeholders — you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them
Nice to Have
- Experience supporting enterprise or federal government customers with high availability requirements
- Experience in a foundational or early SRE hire capacity at a growth stage company
Compensation
The base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location.
Skills
Sre Practices, Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible, Networking, Observability
Similar jobs
DevOps / SRE jobsOwn and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.
Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.