Software Engineer, Site Reliability
Site Reliability Engineer who owns production services end-to-end, writes production code, builds observability and internal tooling, and embeds with product teams to improve reliability and performance. Requires 5+ years experience and strong systems programming skills.
About the job
Responsibilities
- Own critical production services end-to-end, from design and code review through deployment, operation, and incident response
- Profile, benchmark, and rewrite hot paths to eliminate bottlenecks as Hebbia scales
- Lead incident response and drive post-mortem culture, translating findings into code changes and architectural improvements rather than runbooks
- Design and build observability frameworks from scratch, writing custom instrumentation, alerting logic, and debugging tooling
- Define and enforce SLOs across platform services and build the feedback loops that keep engineering teams accountable
- Own capacity planning and cost efficiency: model growth, right-size infrastructure, and write automation that prevents over-provisioning
- Build robust, well-tested internal platforms and deployment tooling held to the same engineering standards as customer-facing code
- Own and continuously improve CI/CD systems
- Embed with product engineering teams as a peer software engineer, contributing directly to production codebases and co-designing systems for reliability
- Partner on infrastructure security through threat modeling, hardening, and automated compliance tooling
Requirements
- 5+ years software development with a track record of writing, shipping, and maintaining production services
- Production-grade proficiency in at least one systems or backend language: Go, Python, C++, or Rust
- Proven experience as a Production Engineer, SRE, or software engineer with a deep infrastructure focus
- Deep understanding of distributed systems
- Container orchestration expertise and hands-on experience debugging complex distributed failures in production
- Working knowledge of OS-level concepts
- Cloud platform fluency (AWS preferred)
- Experience in building and maintaining observability stacks
- Strong CI/CD pipeline expertise and a track record of improving developer velocity without sacrificing safety
Nice-to-Haves
- Background at a company with a Production Engineering or software-focused SRE culture
- Experience building platforms for AI/ML workloads or high-throughput document processing pipelines
Skills
Go, Python, C++, Rust, AWS, Kubernetes, Docker, CI/CD, Observability, Distributed Systems
Similar jobs
DevOps / SRE jobsBuild and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.