Senior Infrastructure Engineer
Senior Infrastructure Engineer designs, builds, and operates cloud infrastructure, developer tooling, observability, and reliability systems at scale, primarily on GCP. Requires high-velocity dev experience, IaC, database scaling, workflow orchestration, and production Python coding.
About the job
Responsibilities
- Design, build, and operate foundational systems powering Scrunch's platform, including cloud infrastructure, developer tooling, observability, and reliability.
Requirements
- Experience in high-velocity software development organizations.
- Experience designing, building, and operating cloud services at scale, preferably on GCP.
- Deployed and operated edge compute functions, understanding edge vs. backend logic.
- Treat infrastructure as software: versioned, tested, deployed via automation (Terraform or similar IaC tools a plus).
- Solid understanding of database performance and scaling patterns; hands-on bottleneck identification and resolution.
- Familiarity with workflow orchestration tools or job queues (Prefect, Airflow, Temporal, Inngest, etc.); design reliable, retryable async pipelines.
- Strong foundation in observable systems: structured logging, distributed tracing, metrics dashboards, alerting.
- Comfortable writing production-quality code (mostly Python) with software engineering mindset.
Nice-to-Haves
- Experience with Terraform or similar IaC tools.
Benefits
- Equity in fast-growing company.
- Medical, dental, vision, life & disability insurance.
- Paid parental leave.
- Home office stipend, phone/internet reimbursement.
- L&D budget.
- Flexible PTO.
- 401(k).
- Team offsites.
Skills
GCP, Terraform, Python, Kubernetes, Prefect, Airflow, Temporal, Postgres, Observability, Distributed Tracing, Structured Logging, Iac, Edge Compute
Similar jobs
DevOps / SRE jobsDesigns, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Own and improve the Linux production infrastructure layer, from performance tuning and incident response to configuration management, orchestration, networking, virtualization, secrets, and observability. The role requires 6+ years of infrastructure or SRE experience and deep Linux expertise.
Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.