Sr. Site Reliability Engineer, tvScientific
Senior Site Reliability Engineer responsible for operating and scaling a cloud-native CTV advertising platform on AWS and Kubernetes. Requires deep Kubernetes and AWS expertise, GitOps with ArgoCD, IaC with Terraform, CI/CD, observability, and incident response experience.
About the job
Responsibilities
- Ensuring the reliability, availability, and performance of production infrastructure and platform services
- Operating and scaling Kubernetes platforms, including governance and support for multi-tenant workloads
- Managing GitOps-based deployment workflows using ArgoCD and Helm
- Driving infrastructure provisioning and change management through Terraform/Terragrunt
- Building and supporting CI/CD automation and deployment workflows using GitHub Actions
- Leading incident response efforts, root cause analysis, and post-incident improvement initiatives
- Reducing operational toil through scripting, tooling, and process automation
- Advancing observability practices across logs, metrics, traces, dashboards, and alerting
- Supporting secure secrets integration, IAM-aware operations, and platform guardrails
- Partnering closely with application, security, and platform teams to improve reliability and delivery outcomes
Requirements
- 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure
- Strong hands-on experience operating AWS in production environments
- Deep expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration
- Proven experience with Kubernetes multi-tenancy, including namespaces, RBAC, quotas, policies, and tenant isolation patterns
- Experience implementing and operating ArgoCD within a GitOps delivery model
- Strong hands-on experience with Helm
- Strong experience with Terraform/Terragrunt for infrastructure provisioning and environment management
- Solid scripting and automation skills using Bash and/or Python
- Experience building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions
- Strong troubleshooting skills across Linux, containers, IAM, networking, and distributed systems
- Experience with monitoring, alerting, and observability in production environments
- Demonstrated ownership mindset with experience handling incidents, resolving production issues, and driving follow-through after outages
- Strong collaboration and communication skills, with the ability to work effectively across engineering, security, and platform teams
- Bachelor’s degree in computer science, engineering, a related field or equivalent experience
- Demonstrated ability to use AI to improve speed and quality in your day-to-day workflow for relevant outputs
- Strong track record of critical evaluation and verification of AI-assisted work (e.g., testing, source-checking, data validation, peer review)
- High integrity and ownership: you protect sensitive data, avoid over-reliance on AI, and remain accountable for final decisions and deliverables
Skills
AWS, Kubernetes, Argo CD, Helm, Terraform, Terragrunt, GitHub Actions, Python, Bash, Linux, IAM, Observability
Similar jobs
DevOps / SRE jobsDesigns, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Own and improve the Linux production infrastructure layer, from performance tuning and incident response to configuration management, orchestration, networking, virtualization, secrets, and observability. The role requires 6+ years of infrastructure or SRE experience and deep Linux expertise.
Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.