# Senior Site Reliability Engineer

**Company:** [Carta](https://hotfix.jobs/companies/carta)
**Location:** London, United Kingdom
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** AWS, GCP, Azure, Kubernetes, Terraform, Ansible, CloudFormation, Python, Docker, Prometheus, Grafana, Datadog, Postgres, gRPC, REST APIs
**Posted:** 2026-08-03

> The Senior Site Reliability Engineer will build and scale cloud infrastructure, observability, networking, and platform services that support Carta’s applications. The role requires strong experience with cloud platforms, infrastructure as code, Kubernetes, monitoring, Python, API services, and reliability practices.

## Job Description

## Responsibilities
- Build and scale internal platform offerings for compute, storage, and networking to ensure application reliability and performance.
- Design and implement monitoring, alerting, and incident-response systems.
- Collaborate with application software engineers to guide scalable, long-term system design.
- Improve systems incrementally as the organization expands globally.

## Requirements
- Extensive experience with cloud services such as AWS, Google Cloud, or Azure, including services such as EC2, S3, RDS, and Lambda.
- Experience with Kubernetes or other container orchestration technologies.
- Proficiency with infrastructure-as-code tools such as Terraform, Ansible, or CloudFormation.
- Experience with networking concepts and tools, including Container Network Interface (CNI) and network policy implementations.
- Knowledge of proxies and service mesh is a plus.
- Strong knowledge of monitoring and observability practices and tools such as Prometheus, Grafana, ELK Stack, or Datadog.
- Proficiency in Python and ability to write efficient, maintainable, scalable code.
- Experience designing, deploying, and maintaining API services, with understanding of RESTful and/or GraphQL API design principles.
- AI fluency, including using AI tools in daily work and building agents to reduce toil.
- Experience operating CI/CD and applying associated best practices is appreciated but not essential.
- Strong communication and collaboration skills.

## Technical Stack
- Python
- Java
- Terraform
- gRPC
- Docker
- Kubernetes
- PostgreSQL
- AWS

## Similar jobs

- [Senior DevSecOps Engineer](https://hotfix.jobs/jobs/9fad0d81-f196-425c-b1d6-cf7adc34ac02) - Shield AI - London, United Kingdom
- [Senior Production Engineer](https://hotfix.jobs/jobs/98eb4848-f8e8-4b98-95bd-c5f2b852c00f) - Clear Street - London, United Kingdom
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Engineer - Platform](https://hotfix.jobs/jobs/df84a55a-dda0-45b3-afb8-9680e622a064) - Hudl - Remote - £66k – £111k/yr
- [Senior Software Engineer, DevOps](https://hotfix.jobs/jobs/bfede940-874a-457b-a7d5-afaa6be82cff) - Muck Rack - Remote - €95k – €110k/yr

**Apply:** https://hotfix.jobs/jobs/ee80c6e1-cdf2-4918-8dee-7392a2b096c6
**Canonical:** https://hotfix.jobs/jobs/ee80c6e1-cdf2-4918-8dee-7392a2b096c6