Staff Software Engineer, Continuous Integration
Build and operate highly reliable continuous integration infrastructure, intelligent test-selection systems, and incident-response automation at scale. The role requires 10+ years of experience with large-scale CI/CD systems, container orchestration, and developer productivity tooling.
About the job
Responsibilities
- Design and build highly reliable, scalable CI infrastructure supporting thousands of daily builds across multiple cloud providers.
- Develop intelligent test-selection systems that reduce CI time while maintaining code quality.
- Build and improve incident-response automation, including cluster load shedding, automatic recovery, and observability tooling.
- Improve test-infrastructure reliability through flake detection, quarantine systems, and test-state management.
Requirements
- 10+ years of relevant industry experience building and operating large-scale CI/CD systems.
- Deep experience with CI orchestration tools such as Buildkite, Jenkins, GitHub Actions, or similar.
- Strong focus on developer productivity and reducing friction in the software development lifecycle.
- Experience with container orchestration at scale.
- Excellent communication skills and ability to support internal partners.
- Strong commitment to reliability and building systems that avoid repeated failure modes.
Nice to haves
- Experience with merge queues and branch management at scale.
- Experience with test infrastructure, including intelligent test selection and flake management.
- Experience building CLI tools and developer-facing services.
- GitHub API and automation experience.
Compensation and benefits
- Annual salary: £325,000–£390,000 GBP.
- Bachelor's degree or equivalent combination of education, training, and experience.
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Office space for collaboration.
- Visa sponsorship may be available depending on the role and candidate.
Skills
Continuous Integration, Continuous Delivery, Buildkite, Jenkins, GitHub Actions, Kubernetes, Github Api, Cli Tools, Test Selection, Test Infrastructure, Merge Queues, Branch Management, Observability, Cloud Infrastructure, Automation
Similar jobs
DevOps / SRE jobsBuild and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.
Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.
Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.