Staff Platform Engineer
Staff Platform Engineer building and scaling the internal DevSecOps platform with IaC, Kubernetes, and multi-cloud automation. Requires 8+ years cloud/infra experience to drive reliability, cost optimization, zero-downtime strategies, and AI-enabled tooling.
About the job
Key Responsibilities
- Design, implement, and manage infrastructure as code (IaC) solutions using tools like Crossplane, Terraform, and Helm Charts to provision and manage cloud resources.
- Collaborate with product managers, developers, InfoSec, and other stakeholders to define platform requirements, scope, and priorities.
- Collaborate with development, operations, and quality assurance teams to streamline the software delivery process and ensure high-quality releases.
- Monitor, troubleshoot, and optimize the performance and cost of our DevOps tools and platforms to ensure reliability, availability, and scalability.
- Design zero-downtime maintenance strategies for fleets of Kubernetes clusters and cloud resources.
- Evaluate industry trends and recommend new tools, technologies, and best practices to improve our DevOps processes and workflows.
- Mentor junior platform engineers, providing guidance, mentorship, and support to foster their growth and development.
- Act as a force multiplier to engineers and development teams by providing expert guidance on integrating with the internal platform, best practices, tools, and methodologies.
- Collaborate with security teams to integrate security controls and best practices into our DevOps processes and infrastructure.
- Document technical designs, procedures, and configurations to ensure knowledge sharing and maintain system integrity.
- Contribute to a culture of innovation, collaboration, and continuous improvement within the TechOps team and across the organization.
- Eliminate toil through the automation of existing processes with modern tooling, including AI tools and agentic workflows.
Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field; or equivalent work experience.
- 8+ years of experience working as a Cloud Engineer, or similar role, with a strong background in software development and infrastructure operations.
- Proficiency with containerization technologies (Docker, Kubernetes) and orchestration tools.
- Proficiency with declarative management tools such as Terraform, Crossplane, or CloudFormation.
- Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Experience in scripting and programming languages such as Go, Python, or Bash.
- Hands-on experience with CI/CD tools such as Jenkins or ArgoCD.
- Exceptional troubleshooting and problem-solving skills, with the ability to quickly diagnose and resolve technical issues.
- Exceptional communication and collaboration skills, with the ability to work effectively in a cross-functional team environment.
- Proven track record of driving process improvements and implementing best practices in DevOps methodologies.
Preferred Qualifications
- Experience with monitoring and logging tools such as Prometheus, Grafana, or OpenTelemetry.
- Knowledge of security best practices and compliance requirements in cloud environments.
- Experience with agile development methodologies and DevOps practices in a fast-paced, dynamic environment.
Skills
Terraform, Crossplane, Kubernetes, Docker, AWS, Azure, GCP, Go, Python, Bash, Jenkins, Argo CD, Helm, Prometheus, Grafana
Similar jobs
DevOps / SRE jobsLeads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting mission-critical payment systems. Requires 8+ years of distributed-systems experience and deep expertise in infrastructure as code, Kubernetes, automation, and cloud networking.
Leads the design and deployment of AI-enabled manufacturing systems, MES, connected-factory infrastructure, and automation for aircraft production. Requires a bachelor’s degree and 8+ years of experience in digital manufacturing, industrial automation, or software-enabled operations.
Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.
Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.
Designs, automates, and operates AWS infrastructure, shared development environments, and container platforms. The role requires strong experience with Kubernetes, infrastructure as code, environment lifecycle automation, cloud security, compliance, and cost optimization.