Build and evolve the infrastructure platform that deploys and operates customer environments. The role focuses on Kubernetes, infrastructure as code, deployment automation, observability, reliability, security, and collaborative continuous delivery practices.
120k – 180k/yr
On-site5+ YOEDevOps / SRE
About the role
How We Work
Practice Extreme Programming and Continuous Delivery.
Use pair programming as the default development approach.
Work in small batches and integrate continuously.
Automate repetitive work and test first whenever practical.
Optimize for learning, feedback, and team outcomes.
Responsibilities
Build and improve the platform that deploys and operates customer environments.
Develop infrastructure as code using Ansible, Terraform, Helm, Zarf, Big Bang, Packer, and related tooling.
Improve Kubernetes platforms and surrounding systems.
Build deployment automation that reduces risk and removes manual work.
Pair with engineers to design, implement, and troubleshoot platform capabilities.
Write automated tests for infrastructure and deployment workflows.
Improve observability, reliability, security, and recoverability.
Write documentation that explains why systems and processes exist.
Decompose large problems into safe, incremental improvements.
Investigate incidents beyond immediate fixes to prevent recurrence.
Requirements
Comfortable working across most of the following areas:
Kubernetes
Linux
Networking fundamentals
Git and trunk-based development
GitLab CI
Infrastructure as code
Ansible
Terraform
Helm
Zarf
Containers
PKI and certificate management
Secrets management
Observability
Infrastructure testing
Troubleshooting distributed systems
Ability to learn quickly and become productive in unfamiliar systems.
Collaborative approach to design, implementation, troubleshooting, and communicating tradeoffs.
Success Measures
Deploy more frequently and safely.
Recover from failures faster.
Remove manual work and reduce operational complexity.
Improve documentation and confidence through testing.
Deliver changes in small, reversible increments.
Make the platform easier to operate and evolve.
Values
Simplicity over cleverness.
Evidence over opinion.
Learning over ego.
Automation over repetition.
Continuous improvement over perfection.
Team outcomes over individual heroics.
Skills
KubernetesLinuxTerraformAnsibleHelmzarfpackergitlab ciGitContainerspkisecrets managementObservabilityDistributed SystemsInfrastructure As Code
Senior Site Reliability Engineer responsible for designing, operating, and improving reliable, scalable production systems. The role requires 5+ years of reliability-focused engineering experience plus expertise in cloud platforms, infrastructure as code, Kubernetes, programming, observability, and critical incident response.
120k – 175k/yrRemote5+ YOEDevOps / SRE
Senior Network Engineer
CommandLinkUnited States
Senior Network Engineer building and supporting carrier interconnects, private circuits, NNIs, and cloud connectivity for a managed network services provider. Requires hands-on service provider experience with Layer 2/3 protocols and direct carrier coordination.
Develops and maintains internal software tools to accelerate engineering workflows for cutting-edge aircraft, integrating EDA, CAD, and PLM systems. Requires 5+ years experience with Python, C++, JavaScript, SQL, Docker, CI/CD, and cloud platforms.
120k – 190k/yrOn-site5+ YOEDevOps / SRE
Senior Infrastructure Engineer
Bland AISan Francisco, CA
Builds and scales distributed systems for real-time voice processing, ML inference, and telephony integration using Kubernetes. Requires 5+ years experience with cloud infrastructure, real-time systems, and tools like Terraform and Datadog.
120k – 200k/yrOn-site5+ YOEDevOps / SRE
Senior Infrastructure Engineer
LiveKitUnited States
Builds and owns foundational infrastructure for globally distributed systems, implements SRE objectives in Golang, manages Kubernetes clusters, and leads incident response. Requires expertise in software engineering, systems administration, and multi-region operations.