Senior or Staff Software Engineer, SRE/Platform Team
This role builds and operates scalable platform infrastructure, automates operational and database workflows, and improves reliability, observability, and deployment practices. It requires at least eight years of platform experience, production systems expertise, Kubernetes, cloud, Linux, and SQL datastore experience.
About the job
Responsibilities
- Identify system bottlenecks and improve database and infrastructure performance.
- Build infrastructure and configuration as code using Kubernetes and Terraform.
- Establish and maintain observability and monitoring systems.
- Define and implement CI/CD best practices and deployment workflows.
- Collaborate with engineering teams to architect scalable, observable services.
- Participate in the on-call rotation, troubleshooting and resolving production incidents.
- Develop software focused on operations, infrastructure, automation, and internal services.
- Automate data center functions and database operations using Kubernetes and custom services.
Requirements
- At least 8 years of platform experience.
- Experience operating reliable production systems at scale.
- Knowledge of Linux systems internals.
- Ability and willingness to automate tasks.
- Experience managing PostgreSQL for high-throughput systems, or comparable SQL datastore experience.
- Operational experience deploying and managing Kubernetes.
- Experience with cloud providers such as AWS, Google Cloud, or Microsoft Azure.
Nice-to-haves
- Recent experience writing Go or Rust.
- Experience with ScyllaDB.
- Experience with Redis, Kafka, etcd, or ClickHouse.
Compensation and benefits
- Senior Software Engineer base salary in the UK: GBP 100,000–GBP 125,000.
- Staff Software Engineer base salary in the UK: GBP 125,000–GBP 145,000.
- Competitive equity program.
- Comprehensive and inclusive benefits.
Skills
Kubernetes, Terraform, Linux, Postgres, AWS, GCP, Microsoft Azure, Go, Rust, Scylladb, Redis, Kafka, Etcd, ClickHouse, CI/CD
Similar jobs
DevOps / SRE jobsLeads production reliability for Grafana Cloud’s multi-tenant database products, partnering with product engineering teams to improve SLOs, scalability, observability, automation, and incident response. Requires 8+ years of engineering experience, including substantial SRE or production engineering work, plus strong Kubernetes and cloud expertise.
Site Reliability Engineers build and operate scalable production infrastructure, automate operational workflows, and improve observability, incident response, and service reliability. The role spans Intermediate through Senior Staff levels and requires experience with Kubernetes, infrastructure as code, cloud platforms, and software engineering.
Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.
Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.