Senior Software Engineer - Observability and Reliability
Build observability tools and platforms (metrics, logging, tracing, alerting) using Go, OpenTelemetry, and Kubernetes. Requires 5+ years experience building high-quality software that other engineers use.
About the job
What You Will Be Doing
- Build observability tools and platforms, including: metrics, logging, distributed tracing, dashboarding, alerting, application performance management
- Build with modern tools and languages like Go, Open Telemetry and Kubernetes
- Participate in on-call rotation and ensure uptime of services
- Create runtime tools/processes that optimize cloud triaging and limit downtime
- Define best practices around making our systems and services measurable
- Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies. We expect successful candidates to be coding a majority of their time
Qualifications We Need
- Strong Computer Science fundamentals
- 5+ years industry experience building and maintaining high-quality software, especially software other engineers use
- You apply a product mindset to infrastructure systems and feel accomplished enabling others
- Desire to be a great teammate and have fun at work
- Strong sense of craftsmanship, and a healthy academic curiosity
Qualifications We Want (also, skills you’ll learn!)
- Experience building systems for data analytics
- Distributed systems monitoring and profiling skills
- Knowledge of cloud application security models
- Administered cloud service infrastructure (GCP, AWS, Azure)
- Startup experience
Compensation & Benefits
- The base salary range for this position is $170k - $240k annually
- Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience
- Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work
- This role is eligible for stock options, as well as a comprehensive benefits package
- Equity
- Generous health benefits
- Flexible time off policy
- Paid bonding time for all new parents
- Traditional and Roth 401k
- Commuter and FSA benefits
- Lunch Program
- Dog friendly office
Skills
Go, OpenTelemetry, Kubernetes, GCP, AWS, Azure, Distributed Systems, Observability, Metrics, Logging, Distributed Tracing, Alerting, Application Performance Management
Similar jobs
DevOps / SRE jobsOwn foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.