Software Engineer - Infrastructure
Infrastructure engineer responsible for maintaining and scaling Kubernetes fleets, improving CI/CD, and making product-level code changes in Python or Go to support autonomous drone platform needs.
Site Reliability Engineer II responsible for designing, deploying, and maintaining multi-cloud infrastructure (Azure primary, AWS/GCP) for Illumio's SaaS products. Focus on IaC, CI/CD pipelines, monitoring, incident response, automation, and improving reliability/scalability in collaboration with engineering and security teams. Requires 2+ years SRE/DevOps experience with Azure.
As an SRE Engineer II, you will be responsible for managing our multi-cloud infrastructure on Azure, AWS and/or GCP. As and when required, you will be responsible for designing new services and applications in the cloud(s) and take them from development to production while working closely with Engineering, SRE/OPS, and Security teams.
On a day-to-day basis, you will work on enhancing system reliability and scalability of Illumio SaaS products, and drive continuous improvement initiatives.
Infrastructure engineer responsible for maintaining and scaling Kubernetes fleets, improving CI/CD, and making product-level code changes in Python or Go to support autonomous drone platform needs.
Technical generalist embeds with operations teams to identify high-impact problems and rapidly builds AI agents, automations, and tools to eliminate friction. Requires 2-5 years software engineering experience with Python/TypeScript proficiency and business impact focus.
Software Engineer building and operating Astronomer's multi-tenant cloud platform and Astro Private Cloud. Focus on Kubernetes, IaC (Terraform), cloud networking (AWS/GCP/Azure), observability, and production reliability with on-call duties. 0-4 years experience; ideal for early-career infra engineers.
Develops Linux-based compute applications for managing virtualization stacks across AI compute servers, integrates with AI hardware like GPUs and NICs, and optimizes performance for AI/ML workloads in datacenters. Requires Linux kernel familiarity, systems programming, and hardware integration skills.
Design, automate, and maintain reliable infrastructure across cloud and colocation environments. Build monitoring, self-service tools, and standards for distributed systems in a Linux-heavy stack.