Senior Site Reliability Engineer
Senior Site Reliability Engineer building and scaling Teleport's secure SaaS cloud infrastructure. Focus on global scaling, observability, automation, incident response, and on-call in a security-first environment using Go and Kubernetes.
About the job
What You'll Do
- Re-engineer the core Teleport product to scale globally and optimize routing latency for distributed teams.
- Re-write portions of the core Teleport product to enable cloud product goals.
- Build out monitoring and observability stack to alert on production issues and minimize false positives.
- Work on automation to eliminate highest toil activities.
- Execute on traditional operational challenges such as patching, scaling, backup and restore, disaster recovery.
- Investigate outages and incidents experienced by customers.
- Participate in on-call rotation for 24/7/365 system uptime.
Requirements
- Strong experience in Linux systems, networking, containers, and troubleshooting.
- Solid Go and Kubernetes development experience.
- Strong experience developing scripts, automation, or lightweight programs; submitting patches to product codebase; or building tooling that incorporates AI agents into operational workflows.
- Experience operating in a team where sound security choices are critical and reasoning about correctness and system invariants (e.g. formal or property-based methods) is valued.
- Intellectual curiosity and willingness to master new technologies.
- Transparency, honesty, and a no-ego mindset.
- Excellent communication skills.
- Willingness to collaboratively work with Teleport’s engineers on a coding challenge in Go as part of the interview process.
Nice-to-Haves
- AWS Cloud experience (preferred); GCP experience acceptable.
- Experience with systems observability tools: Prometheus, Grafana, Loki, etc.
- Operate and support the observability platform to maintain visibility and reliability.
Skills
Go, Kubernetes, Linux, Networking, Containers, AWS, Prometheus, Grafana, Loki, Observability, Automation
Similar jobs
DevOps / SRE jobsOwn and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.
Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.