Staff Platform Engineer
Owns the technical direction of an AI-first platform, defining architecture, standards, and self-service capabilities that help engineering teams ship securely and efficiently. The role requires multi-team technical leadership, strong judgment on foundational trade-offs, and the ability to connect platform strategy to business outcomes.
About the job
Responsibilities
- Establish a coherent 12–18 month technical direction for the platform, including trade-offs, milestones, and alignment across engineering leadership.
- Drive broad adoption of platform capabilities through thoughtful design, documentation, and reduced friction for provisioning and deployment.
- Define architectural patterns and standards that enable autonomous and AI-augmented workflows across the organisation.
- Establish platform standards for security, reliability, and operational excellence that teams adopt through ease of use rather than gatekeeping.
- Reduce misconfiguration, security exceptions, operational incidents, and platform bottlenecks across the engineering estate.
- Serve as a technical authority and strengthen senior engineers through technical challenge, feedback, and exposure to complex problems.
- Make high-judgement decisions on foundational trade-offs such as build versus buy, depth versus breadth, and short-term pragmatism versus long-term coherence.
- Navigate disagreement among senior engineers, architects, and stakeholders to reach durable technical decisions.
- Design systems, standards, and enablement approaches that multiply the effectiveness of engineers across multiple teams.
- Define how AI and automation are embedded into the platform at an architectural level.
- Shape the platform roadmap using internal-customer feedback, usage data, and analysis of self-service friction.
- Identify architectural, security, and operational risks before they become incidents.
- Connect platform decisions to business outcomes and communicate platform strategy to senior leadership.
- Develop senior engineers and model the organisation's engineering culture.
Requirements
- Experience defining architecture, standards, and patterns for a platform domain over a multi-quarter horizon.
- Experience making and communicating foundational technical trade-offs.
- Experience influencing senior engineers, architects, stakeholders, and teams without directly delivering every initiative.
- Experience embedding AI and automation into platform architecture and engineering workflows.
- Understanding of platform adoption dynamics, internal customer needs, and barriers to self-service.
- Ability to identify and influence security, reliability, operational, and cost-efficiency risks across an engineering organisation.
- Ability to translate engineering constraints and platform strategy into clear business language.
- Experience developing senior engineers through technical challenge and feedback.
Compensation and Benefits
- Equity offered across all roles.
- £5,000 training and conference budget for individual and group development.
- 25 days of holiday plus 8 bank holidays (33 days total).
- Company pension scheme via Penfold.
- Mental health support and therapy via Spectrum.life.
- Individual wellbeing allowance via Juno.
- Private healthcare insurance through AXA.
- Income protection.
Skills
Platform Engineering, AI, Automation, Self-Service Infrastructure, Cloud Infrastructure, Architecture, Security, Reliability Engineering, Operational Excellence, Infrastructure As Code, Developer Productivity, Technical Standards
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.
Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.
Staff Engineer responsible for deploying, integrating, maintaining, and developing an AI training factory across isolated environments. The role requires 7+ years of related experience, cloud and Kubernetes expertise, Linux networking knowledge, application support skills, and automation experience.