Senior Site Reliability Engineer
The Senior Site Reliability Engineer will build and scale cloud infrastructure, observability, networking, and platform services that support Carta’s applications. The role requires strong experience with cloud platforms, infrastructure as code, Kubernetes, monitoring, Python, API services, and reliability practices.
About the job
Responsibilities
- Build and scale internal platform offerings for compute, storage, and networking to ensure application reliability and performance.
- Design and implement monitoring, alerting, and incident-response systems.
- Collaborate with application software engineers to guide scalable, long-term system design.
- Improve systems incrementally as the organization expands globally.
Requirements
- Extensive experience with cloud services such as AWS, Google Cloud, or Azure, including services such as EC2, S3, RDS, and Lambda.
- Experience with Kubernetes or other container orchestration technologies.
- Proficiency with infrastructure-as-code tools such as Terraform, Ansible, or CloudFormation.
- Experience with networking concepts and tools, including Container Network Interface (CNI) and network policy implementations.
- Knowledge of proxies and service mesh is a plus.
- Strong knowledge of monitoring and observability practices and tools such as Prometheus, Grafana, ELK Stack, or Datadog.
- Proficiency in Python and ability to write efficient, maintainable, scalable code.
- Experience designing, deploying, and maintaining API services, with understanding of RESTful and/or GraphQL API design principles.
- AI fluency, including using AI tools in daily work and building agents to reduce toil.
- Experience operating CI/CD and applying associated best practices is appreciated but not essential.
- Strong communication and collaboration skills.
Technical Stack
- Python
- Java
- Terraform
- gRPC
- Docker
- Kubernetes
- PostgreSQL
- AWS
Skills
AWS, GCP, Azure, Kubernetes, Terraform, Ansible, CloudFormation, Python, Docker, Prometheus, Grafana, Datadog, Postgres, gRPC, REST APIs
Similar jobs
DevOps / SRE jobsDesigns and operates secure development infrastructure and CI/CD pipelines for autonomous defence systems. Requires at least five years of DevOps or related experience, plus UK defence or regulated national-security experience and knowledge of Secure by Design and assurance practices.
Own production reliability and operational excellence by supporting incidents while building automation, observability, self-healing, and diagnostic tooling. The role requires strong Python, cloud-native, Kubernetes, distributed-systems, and infrastructure-as-code experience.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Senior Platform Engineer responsible for architecting scalable, secure infrastructure and improving reliability, observability, and production operations. The role requires strong AWS, Infrastructure as Code, and Kubernetes experience, along with technical leadership and mentoring skills.
Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.