Sr Software Engineer, Infrastructure
Senior Software Engineer builds and automates scalable AWS infrastructure, manages Kubernetes clusters, and implements observability frameworks. Requires 5+ years Python experience, IaC expertise, and strong cloud/DevOps skills.
About the job
Responsibilities
- Architect and automate production-grade infrastructure on AWS using Terraform or Pulumi.
- Manage and scale containerized workloads using AKS (Azure Kubernetes Service) or EKS, focusing on cluster security and resource efficiency.
- Architect robust deployment pipelines using GitHub Actions, managing both GitHub-hosted and self-hosted runners.
- Create infrastructure for "Observable by Default" frameworks ensuring new applications are secure with logging and metrics enabled.
- Build internal CLI tools, AI plugins, and automation scripts to streamline developer workflows.
- Collaborate cross-functionally with Security, Engineering, Infrastructure, and Support teams.
- Mentor junior engineers, participate in code reviews, and document solutions and failure triage playbooks.
Requirements
- 5+ years production-level experience with strong proficiency in Python (required).
- Expert-level Terraform (modules, state management) or Pulumi (preferred).
- Hands-on experience with AWS (or Azure/GCP), Kubernetes, Docker.
- Experience building/troubleshooting integrations between infrastructure, data pipelines, and observability platforms.
- Advanced knowledge of GitHub Actions, GitHub Runners.
- Strong observability mindset: logging, metrics, tracing; experience with Datadog, Prometheus, or ELK.
- Proficiency in distributed systems concepts like Kafka or messaging queues.
- Ability to operate independently on ambiguous projects.
Skills
Python, Terraform, Pulumi, AWS, Kubernetes, Docker, GitHub Actions, Datadog, Prometheus, Elk, Kafka
Similar jobs
DevOps / SRE jobsBuild and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Own and improve the Linux production infrastructure layer, from performance tuning and incident response to configuration management, orchestration, networking, virtualization, secrets, and observability. The role requires 6+ years of infrastructure or SRE experience and deep Linux expertise.
Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.