Skip to content
Scale AIScale AI

Infrastructure Software Engineer, Apps Platform

Build and operate cloud-agnostic deployment and observability infrastructure across public clouds and on-premises environments. The role requires 5+ years of infrastructure experience, strong networking and IaC expertise, and ownership of production systems and cross-functional projects.

About the job

Responsibilities

  • Build and evolve deployment and observability layers across cloud providers and on-premises environments.
  • Expand platform deployment coverage across on-premises infrastructure and major cloud platforms.
  • Ensure fast, secure, and reproducible deployments across supported cloud service providers and on-premises environments.
  • Improve platform observability so infrastructure teams can operate customer deployments efficiently.
  • Partner with internal platform-deployment teams to understand needs, debug issues, and build supporting tooling.
  • Respond to incidents and production issues, conduct root-cause analysis, and implement preventive fixes.
  • Help develop and maintain the deployment and observability product roadmap.
  • Lead architecture reviews and own projects end to end, from design through deployment.

Requirements

  • 5+ years of experience building and deploying enterprise and public-sector solutions across AWS, Azure, Google Cloud, Oracle Cloud Infrastructure, and on-premises environments.
  • Expertise in infrastructure design, networking, VPNs, load balancers, and firewalls.
  • Proficiency with infrastructure as code, containerization, and orchestration.
  • Comprehensive understanding of CI/CD pipelines and software delivery principles, using GitHub Actions or CircleCI.
  • Experience with GitOps-style deployments and consistent, repeatable deployment practices.
  • Strong debugging skills and ability to navigate performance and security trade-offs in production systems.
  • Ability to context-switch between incident response and proactive product development.

Nice to Have

  • Founder or early-engineer experience at an infrastructure-focused startup, owning a product end to end.
  • Experience running secure workloads in multi-tenant or untrusted environments, such as FaaS, CI sandboxes, or remote notebooks.
  • Open-source contributions to systems or developer-tools projects.
  • Production on-call and incident-response experience.

Skills

AWS, Azure, GCP, Oracle Cloud Infrastructure, Kubernetes, Terraform, Docker, CI/CD, GitHub Actions, CircleCI, GitOps, Networking, Vpns, Load Balancers, Firewalls

Cloudflare

Cloudflare

London, United Kingdom

Software Engineer: Resiliency - Deploy at Scale
No salary listedHybrid4+ YOEDevOps / SRE

Build and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

GitLab

GitLab

United Kingdom

Site Reliability Engineer, Infrastructure Platforms
No salary listedRemote5+ YOEDevOps / SRE

Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Writer

Writer

London, United Kingdom

Infrastructure Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.