# Senior Software Engineer, Infrastructure

**Company:** [Voltus](https://hotfix.jobs/companies/voltus)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $160k – $190k/yr
**Experience:** 6+ years
**Skills:** Go, Python, Kubernetes, AWS, Terraform, Docker, Nomad, Consul, Vault, Prometheus, Grafana, OpenTelemetry, GitHub, Buildkite, Argo Cd
**Posted:** 2026-08-26

> Senior software engineer responsible for operating and evolving Voltus’s infrastructure platform across AWS, Kubernetes, Nomad, observability, stateful systems, and developer tooling. The role requires 6+ years of engineering experience, deep production Kubernetes and AWS expertise, and strong Go or Python skills.

## Job Description

## Responsibilities
- Own core platform services and major migrations end to end, from proposal through production, including stateful and business-critical systems with minimal customer-visible downtime.
- Operate containerized workloads and delivery systems, including orchestration, scheduling, GitOps deployments, progressive rollouts, rollback, and service mesh.
- Bring production Kubernetes expertise to a team currently running Nomad and help determine the appropriate use of each scheduler.
- Architect and operate AWS infrastructure, including multi-account governance, IAM, cross-account access, VPC, Transit Gateway, PrivateLink, DNS, certificates, and egress paths.
- Manage identity, secrets, encryption, workload identity, least-privilege access, SSO/OIDC, machine-to-machine credentials, Vault, and key rotation.
- Operate databases, replication, message brokers, caches, time-series stores, backups, and tested restores.
- Build observability through distributed tracing, metrics, SLOs, dashboards, alerts, and testing frameworks; consolidate monitoring into infrastructure defined as code.
- Build infrastructure as code and developer tooling using Terraform, GitHub, Buildkite, Docker, Nomad, and internal tools, including importing manually created infrastructure and detecting drift.
- Build secure infrastructure and guardrails for AI-assisted development and access to internal systems.
- Participate in on-call ownership and plan operational changes around live dispatch windows with safe rollback paths.
- Read and improve unfamiliar systems, document dependencies, mentor engineers, and coordinate cross-team infrastructure work.

## Requirements
- 6+ years of professional software engineering experience, including several years in DevOps or SRE operating production systems.
- Strong production software development experience in Go and/or Python, including services, tooling, tests, and maintainable code.
- Experience owning infrastructure projects end to end, preferably migrations of stateful or business-critical systems.
- Deep AWS experience with multi-account organizations, IAM, cross-account access, VPC and network design, DNS, secrets management, and encryption key management.
- Deep hands-on production Kubernetes experience, including cluster upgrades, networking and ingress, RBAC, resource management, autoscaling, and debugging workloads under load.
- Strong infrastructure-as-code skills with Terraform or a similar tool.
- Strong monitoring and observability experience, including metrics, alerts, dashboards, distributed tracing, and open-source tooling such as Prometheus, Grafana, Elasticsearch, or OpenSearch.
- Hands-on experience operating stateful systems such as relational databases and message brokers, including upgrades, replication changes, or restores.
- Clear communication, documentation, mentoring, and cross-team coordination skills.
- Interest in or experience with AI-assisted development tools such as Claude Code, MCP, or agents.

## Nice to Have
- Experience with Nomad, Consul, and Vault.
- Experience introducing Kubernetes to an organization or operating it alongside another scheduler.
- Experience running self-hosted CI and build tooling, including Jenkins, Argo CD, or artifact repositories.
- Experience with OpenTelemetry or consolidating overlapping monitoring tools.
- Experience testing event-driven workflows and using managed streaming platforms such as Amazon MSK.
- Experience with AWS Organizations, Control Tower, service control policies, or leading infrastructure governance.

## Compensation
- Annual salary: $160,000–$190,000.

## Similar jobs

- [Senior Software Engineer, Infrastructure](https://hotfix.jobs/jobs/46c3fea2-64c0-48d6-aef4-3cd59872c8e8) - VSCO - San Francisco, CA - $165k – $185k/yr
- [Senior Software Engineer](https://hotfix.jobs/jobs/863cf815-4e44-4e0a-a0a4-cb4608fafeab) - Imply - Remote - $155k – $215k/yr
- [Senior Software Engineer, Infrastructure](https://hotfix.jobs/jobs/dcbafc48-83a4-41c8-83fa-cbff032ea845) - Commure - Mountain View, CA - $170k – $220k/yr
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Infrastructure Engineer](https://hotfix.jobs/jobs/8a829a2b-c825-45b8-ac7e-8a92504dfad2) - Gumloop - San Francisco, CA - $150k – $300k/yr

**Apply:** https://hotfix.jobs/jobs/72965952-2743-4317-b616-77292db199b2
**Canonical:** https://hotfix.jobs/jobs/72965952-2743-4317-b616-77292db199b2