Senior Software Engineer, Infrastructure
Senior software engineer responsible for operating and evolving Voltus’s infrastructure platform across AWS, Kubernetes, Nomad, observability, stateful systems, and developer tooling. The role requires 6+ years of engineering experience, deep production Kubernetes and AWS expertise, and strong Go or Python skills.
About the job
Responsibilities
- Own core platform services and major migrations end to end, from proposal through production, including stateful and business-critical systems with minimal customer-visible downtime.
- Operate containerized workloads and delivery systems, including orchestration, scheduling, GitOps deployments, progressive rollouts, rollback, and service mesh.
- Bring production Kubernetes expertise to a team currently running Nomad and help determine the appropriate use of each scheduler.
- Architect and operate AWS infrastructure, including multi-account governance, IAM, cross-account access, VPC, Transit Gateway, PrivateLink, DNS, certificates, and egress paths.
- Manage identity, secrets, encryption, workload identity, least-privilege access, SSO/OIDC, machine-to-machine credentials, Vault, and key rotation.
- Operate databases, replication, message brokers, caches, time-series stores, backups, and tested restores.
- Build observability through distributed tracing, metrics, SLOs, dashboards, alerts, and testing frameworks; consolidate monitoring into infrastructure defined as code.
- Build infrastructure as code and developer tooling using Terraform, GitHub, Buildkite, Docker, Nomad, and internal tools, including importing manually created infrastructure and detecting drift.
- Build secure infrastructure and guardrails for AI-assisted development and access to internal systems.
- Participate in on-call ownership and plan operational changes around live dispatch windows with safe rollback paths.
- Read and improve unfamiliar systems, document dependencies, mentor engineers, and coordinate cross-team infrastructure work.
Requirements
- 6+ years of professional software engineering experience, including several years in DevOps or SRE operating production systems.
- Strong production software development experience in Go and/or Python, including services, tooling, tests, and maintainable code.
- Experience owning infrastructure projects end to end, preferably migrations of stateful or business-critical systems.
- Deep AWS experience with multi-account organizations, IAM, cross-account access, VPC and network design, DNS, secrets management, and encryption key management.
- Deep hands-on production Kubernetes experience, including cluster upgrades, networking and ingress, RBAC, resource management, autoscaling, and debugging workloads under load.
- Strong infrastructure-as-code skills with Terraform or a similar tool.
- Strong monitoring and observability experience, including metrics, alerts, dashboards, distributed tracing, and open-source tooling such as Prometheus, Grafana, Elasticsearch, or OpenSearch.
- Hands-on experience operating stateful systems such as relational databases and message brokers, including upgrades, replication changes, or restores.
- Clear communication, documentation, mentoring, and cross-team coordination skills.
- Interest in or experience with AI-assisted development tools such as Claude Code, MCP, or agents.
Nice to Have
- Experience with Nomad, Consul, and Vault.
- Experience introducing Kubernetes to an organization or operating it alongside another scheduler.
- Experience running self-hosted CI and build tooling, including Jenkins, Argo CD, or artifact repositories.
- Experience with OpenTelemetry or consolidating overlapping monitoring tools.
- Experience testing event-driven workflows and using managed streaming platforms such as Amazon MSK.
- Experience with AWS Organizations, Control Tower, service control policies, or leading infrastructure governance.
Compensation
- Annual salary: $160,000–$190,000.
Skills
Go, Python, Kubernetes, AWS, Terraform, Docker, Nomad, Consul, Vault, Prometheus, Grafana, OpenTelemetry, GitHub, Buildkite, Argo Cd
Similar jobs
DevOps / SRE jobsOwn and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.
Build and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.
Own foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.