Software Engineer, Platform
Build and operate high-throughput infrastructure for AI engineering workloads, including sandbox runtimes, schedulers, networking, and reliability systems. The role requires deep production infrastructure experience, systems programming expertise, and strong knowledge of orchestration and isolation.
About the job
Responsibilities
- Design, build, and maintain systems that enable Lovable’s AI product.
- Design and operate a gVisor-based runtime environment for agentic workloads.
- Build a high-throughput sandbox scheduler across multiple cloud providers.
- Harden infrastructure against failures, downtime, and slowdowns.
- Ensure infrastructure scales without becoming a bottleneck.
- Plan and implement network infrastructure and cloud strategy.
- Identify and drive reliability improvements across engineering teams.
- Design, ship, and own critical services and systems.
Requirements
- Deep experience building and operating production infrastructure as a software engineer, systems engineer, site reliability engineer, or similar.
- Ability to write clean, performant code and understand systems from the API layer through the runtime.
- Familiarity with container orchestration and sandboxing beyond configuration-level usage.
- Understanding of schedulers, runtimes, and isolation boundaries.
- Strong proficiency in at least one systems-oriented language such as Go, Rust, or C++.
- Track record designing services that handle high throughput and unpredictable load at a global technology company or scale-up.
- Comfort navigating ambiguity and solving problems as they arise.
- Ability to balance security, stability, and speed.
- Based in Stockholm or willing to relocate.
Technical Stack
- Frontend: React, TypeScript
- Backend: Golang, Rust
- Cloud: Cloudflare, Google Cloud, AWS
- Data: ClickHouse, Firestore, Spanner, BigQuery
- DevOps and tooling: CI/CD, OpenTelemetry, Kubernetes, Terraform
Skills
Go, Rust, C++, Kubernetes, Gvisor, Terraform, AWS, GCP, Cloudflare, React, TypeScript, ClickHouse, Firestore, Spanner, BigQuery
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.
Designs and operates secure, highly available cloud infrastructure supporting engineering teams, with a focus on GCP, GKE, Terraform, Kubernetes, observability, and developer self-service. Requires 5–8 years of production infrastructure experience and strong cloud, automation, and Linux expertise.