Skip to content
Shield AIShield AISan Diego, CA

Senior Engineer, Platform Infrastructure

Build and evolve the infrastructure platform that deploys and operates customer environments. The role focuses on Kubernetes, infrastructure as code, deployment automation, observability, reliability, security, and collaborative continuous delivery practices.

120k – 180k/yr
On-site5+ YOEDevOps / SRE

About the role

How We Work

  • Practice Extreme Programming and Continuous Delivery.
  • Use pair programming as the default development approach.
  • Work in small batches and integrate continuously.
  • Automate repetitive work and test first whenever practical.
  • Optimize for learning, feedback, and team outcomes.

Responsibilities

  • Build and improve the platform that deploys and operates customer environments.
  • Develop infrastructure as code using Ansible, Terraform, Helm, Zarf, Big Bang, Packer, and related tooling.
  • Improve Kubernetes platforms and surrounding systems.
  • Build deployment automation that reduces risk and removes manual work.
  • Pair with engineers to design, implement, and troubleshoot platform capabilities.
  • Write automated tests for infrastructure and deployment workflows.
  • Improve observability, reliability, security, and recoverability.
  • Write documentation that explains why systems and processes exist.
  • Decompose large problems into safe, incremental improvements.
  • Investigate incidents beyond immediate fixes to prevent recurrence.

Requirements

  • Comfortable working across most of the following areas:
    • Kubernetes
    • Linux
    • Networking fundamentals
    • Git and trunk-based development
    • GitLab CI
    • Infrastructure as code
    • Ansible
    • Terraform
    • Helm
    • Zarf
    • Containers
    • PKI and certificate management
    • Secrets management
    • Observability
    • Infrastructure testing
    • Troubleshooting distributed systems
  • Ability to learn quickly and become productive in unfamiliar systems.
  • Collaborative approach to design, implementation, troubleshooting, and communicating tradeoffs.

Success Measures

  • Deploy more frequently and safely.
  • Recover from failures faster.
  • Remove manual work and reduce operational complexity.
  • Improve documentation and confidence through testing.
  • Deliver changes in small, reversible increments.
  • Make the platform easier to operate and evolve.

Values

  • Simplicity over cleverness.
  • Evidence over opinion.
  • Learning over ego.
  • Automation over repetition.
  • Continuous improvement over perfection.
  • Team outcomes over individual heroics.

Skills

KubernetesLinuxTerraformAnsibleHelmzarfpackergitlab ciGitContainerspkisecrets managementObservabilityDistributed SystemsInfrastructure As Code

Similar roles

DevOps / SRE jobs
PrizePicks

Senior Site Reliability Engineer

PrizePicksUnited States

Senior Site Reliability Engineer responsible for designing, operating, and improving reliable, scalable production systems. The role requires 5+ years of reliability-focused engineering experience plus expertise in cloud platforms, infrastructure as code, Kubernetes, programming, observability, and critical incident response.

120k – 175k/yrRemote5+ YOEDevOps / SRE
CommandLink

Senior Network Engineer

CommandLinkUnited States

Senior Network Engineer building and supporting carrier interconnects, private circuits, NNIs, and cloud connectivity for a managed network services provider. Requires hands-on service provider experience with Layer 2/3 protocols and direct carrier coordination.

120k – 160k/yrRemote5+ YOEDevOps / SRE
Shield AI

Senior Engineer, Software Engineering Tools (R4913)

Shield AIDallas, TX

Develops and maintains internal software tools to accelerate engineering workflows for cutting-edge aircraft, integrating EDA, CAD, and PLM systems. Requires 5+ years experience with Python, C++, JavaScript, SQL, Docker, CI/CD, and cloud platforms.

120k – 190k/yrOn-site5+ YOEDevOps / SRE
Bland AI

Senior Infrastructure Engineer

Bland AISan Francisco, CA

Builds and scales distributed systems for real-time voice processing, ML inference, and telephony integration using Kubernetes. Requires 5+ years experience with cloud infrastructure, real-time systems, and tools like Terraform and Datadog.

120k – 200k/yrOn-site5+ YOEDevOps / SRE
LiveKit

Senior Infrastructure Engineer

LiveKitUnited States

Builds and owns foundational infrastructure for globally distributed systems, implements SRE objectives in Golang, manages Kubernetes clusters, and leads incident response. Requires expertise in software engineering, systems administration, and multi-region operations.

120k – 250k/yrRemoteDevOps / SRE