Senior Platform Engineer
Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
About the job
Responsibilities
- Design, build, and operate scalable platform infrastructure across Azure, AWS, and private cloud environments.
- Develop reusable infrastructure-as-code modules, automation frameworks, and self-service platform capabilities.
- Own platform initiatives from technical design through production operations and continuous improvement.
- Build and maintain container and Kubernetes platforms, including deployment automation, configuration management, observability, and lifecycle management.
- Develop platform tooling and automation using Terraform, Ansible, Python, and Go.
- Build standardized paved paths for infrastructure provisioning, application deployment, and shared platform services.
- Establish platform standards for reliability, scalability, security, performance, and cost efficiency.
- Support capacity planning, performance tuning, vulnerability remediation, disaster recovery, and platform lifecycle management.
- Build and improve CI/CD pipelines, monitoring, logging, alerting, and developer-facing platform services.
- Troubleshoot complex platform, infrastructure, and application integration issues and lead root-cause analysis.
- Maintain architecture diagrams, technical documentation, operational procedures, and reusable implementation patterns.
- Evaluate technologies and recommend improvements to automation, reliability, and developer productivity.
- Provide technical guidance and contribute to platform architecture, standards, and roadmap decisions.
- Participate in an on-call rotation and scheduled after-hours maintenance.
Requirements
- 7+ years of experience in platform engineering, cloud infrastructure, DevOps, Site Reliability Engineering, or a related discipline.
- Experience designing and operating production infrastructure in Azure, AWS, or comparable public cloud environments.
- Strong experience with reusable infrastructure-as-code modules and automation using Terraform and Ansible.
- Hands-on experience deploying and operating containerized workloads and Kubernetes platforms.
- Proficiency in Python, Go, or another programming language used for platform tooling and automation.
- Experience building CI/CD pipelines, self-service infrastructure, or developer-facing platform services.
- Strong Linux systems administration experience, including deployment, configuration, troubleshooting, patching, and performance analysis.
- Knowledge of cloud and enterprise networking, including VPCs/VNets, subnets, routing, VPNs, load balancing, DNS, and firewalls.
- Experience implementing monitoring, logging, alerting, and observability for production platforms.
- Understanding of platform security, identity and access management, vulnerability remediation, and secure configuration practices.
- Ability to lead complex technical initiatives from design through production deployment.
- Strong technical documentation, communication, collaboration, and organizational skills.
- Bachelor's degree in computer science or a related field, or equivalent practical experience.
Nice-to-haves
- Experience supporting commercial, government, regulated, classified, or air-gapped cloud environments.
- Experience with Microsoft Azure and AWS, including Azure Government or AWS GovCloud.
- Experience with GitOps, Helm, policy as code, secrets management, service catalogs, or internal developer portals.
- Experience designing or supporting internal developer platforms and standardized paved-path tooling.
- Knowledge of private cloud and virtualization platforms such as VMware, Hyper-V, or KVM.
- Experience defining service-level indicators, service-level objectives, incident response practices, and reliability improvements.
- Experience supporting hybrid-cloud architectures across public cloud, private cloud, and on-premises infrastructure.
- Experience in aerospace, defense, manufacturing, or other regulated environments.
- Relevant cloud, Kubernetes, Terraform, or platform-engineering certifications.
Skills
Azure, AWS, Terraform, Ansible, Kubernetes, Python, Go, Linux, CI/CD, Observability, Cloud Networking, Identity And Access Management, Helm, GitOps, VMware
Similar jobs
DevOps / SRE jobsDesigns, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Own and improve the Linux production infrastructure layer, from performance tuning and incident response to configuration management, orchestration, networking, virtualization, secrets, and observability. The role requires 6+ years of infrastructure or SRE experience and deep Linux expertise.
Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.
The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.