Staff Software Engineer - Cloud Infrastructure and Applications
Designs and implements scalable cloud infrastructure for healthcare AI platform using Kubernetes, Terraform, and AWS/GCP. Owns DevOps pipelines, automation, reliability, and security with 8+ years experience.
About the job
What You’ll Do
- Create, implement, and support DevOps strategies and continuous delivery pipelines with cross-functional agile teams
- Own defining and implementing infrastructure, tools, and processes for continuous delivery of change; identify potential issues
- Play a vital role in maturing continuous delivery processes for high availability and quality
- Design and build infrastructure to support existing and upcoming products
- Plan for infrastructure maintainability and foresee weaknesses
- Identify new technologies to improve automation
- Document infrastructure setup and best practices
What We’re Looking For
- 8+ years of relevant work experience
- Hands-on SaaS delivery experience with AWS or GCP systems including incident response
- Experience in automating build, test, package, release, and configuration management
- Experience with Terraform
- Good understanding of Linux/Unix fundamentals and debugging skills
- Strong scripting skills (Bash, Python, NodeJS, Go)
- Experience defining and deploying monitoring, metrics, and logging systems
- Recent hands-on experience creating and managing containerized deployments (Kubernetes)
- Demonstrable experience with networks, security, load balancers, DNS, etc
- Rigor in high-code quality, automated testing, and other engineering best practices
Skills
Terraform, Kubernetes, AWS, GCP, Linux, Python, Bash, Go, Node.js, DevOps, CI/CD, Monitoring, Networking, Security
Similar jobs
DevOps / SRE jobsStaff DevSecOps Engineer designing and automating security controls across AWS infrastructure, containers, CI/CD, and platform services. Requires 7+ years of related experience plus expertise in cloud security, infrastructure as code, hardened images, vulnerability scanning, identity, and secrets management.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.