Staff Platform Engineer - Infra + DevOps
Seasoned Platform Engineer designing and maintaining scalable distributed systems and infrastructure on AWS. Builds foundational patterns, IaC, CI/CD pipelines, and observability for Python services using Kubernetes and serverless. Requires 6+ years of platform engineering experience.
About the job
Responsibilities
- Design and implement foundational patterns and libraries for Python applications, across a range of technologies from API services to event processing
- Utilize Infrastructure as Code (IaC) tools to ensure reproducible and scalable environment setups
- Design and implement infrastructure for applications hosted on AWS, supporting event-driven systems, containerized services on Kubernetes, and serverless functions
- Develop and maintain robust CI/CD pipelines using tools such as Jenkins, ArgoCD
- Automate the lifecycle management of code from development through production, including code promotion and configuration management
- Instrument observability through tools such as CloudWatch and DataDog to monitor and optimize application performance across multiple environments
- Scale infrastructure to meet increasing demand while managing cost effectively
- Define, instrument and measure standards for quality, security, scalability, and availability with a focus on delivering business value
- Deliver turn-key developer experience for local development
- Mentor and develop talent
Requirements
- At least 6 years of professional experience in a platform engineering team within a product-centric company
- Experience designing, implementing, and maintaining scalable distributed systems & infrastructure
- Deep practical knowledge of cloud platforms, advanced software design patterns & architecture, operations and automation, and containerization technologies like Kubernetes
- Experience working with AWS, event-driven systems, containerized services on Kubernetes, and serverless functions
- Experience with Infrastructure as Code (IaC) tools
- Experience with CI/CD pipelines using Jenkins, ArgoCD
- Experience with observability tools such as CloudWatch and DataDog
- Strong written and verbal communication skills
- Strategic thinker with a product-first approach and customer obsession
Nice-to-Haves
- Experience working with an RPC architecture
- Experience at a technology startup and familiar with the challenges of evolving platform maturity
- First hand experience navigating multiple distributed architecture patterns
Skills
Python, AWS, Kubernetes, Infrastructure As Code, Jenkins, Argo CD, CI/CD, CloudWatch, Datadog, Serverless
Similar jobs
DevOps / SRE jobsOwn reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.