Senior Software Engineer, Site Reliability
Senior SRE engineer builds tooling and automation to enhance production system reliability, monitoring microservices, Kubernetes, and ML platforms. Requires 6+ years in software/SRE/DevOps, proficiency in Python/Go, IaC, and observability tools.
About the job
How you’ll make an impact
- Embody and share SRE principles at Upstart
- Exercise state-of-the-art SRE practices throughout the company
- Uphold a culture of visibility, ownership, and responsibility around service reliability
- Implement standards for monitoring microservices, web apps, mobile apps, databases, Kubernetes clusters, and machine learning platforms, in a fast-paced environment
- Improve incident response practices, both within SRE and throughout the company
- Automate away toil that make sense to be automated
What we’re looking for
Minimum requirements:
- Minimum of 6 years combined experience between Software Engineering, Site Reliability, and/or DevOps Engineering including CI/CD, TDD, internal tooling, observability, and other agile development practices
- Proficiency coding Python, Go, JavaScript/TypeScript
- Proficiency with Infrastructure as Code (Terraform, CDK, Cloudformation, etc.)
- Software engineering background with experience building internal tooling from scratch, and other agile development techniques
- Strong software design & architecture skills
- Fundamentally sound with data structures & algorithms
- Experience with on-call and incident management environments
- Experience with observability, monitoring, and reporting tools (e.g., Datadog, Sumologic, etc.)
- Experience supporting SaaS software in a microservice-oriented cloud environment
- Ability to work with multiple teams for enterprise-wide deliverables
- Data/metrics-driven mindset
Preferred qualifications:
- Experience with service mesh
- Full Stack development skills
- Experience building tooling for an observability platform
- Experience leveraging LLM/GenAI to improve SRE efficiency and processes
Skills
Ruby on Rails, React, AWS, Docker, GitHub Actions, Distributed Systems, Service-Oriented Architecture, CI/CD, Infrastructure As Code, A/B Testing
Similar jobs
DevOps / SRE jobsOwn and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.
Own foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.