Senior Software Engineer, Infrastructure
Senior Infrastructure Engineer responsible for building and operating platform primitives including Kubernetes, CI/CD, observability, and developer tooling at a high-growth AI and data platform company.
About the job
Responsibilities
- Steward core platform services: Implement container orchestration, service mesh, ingress, and secrets management at scale.
- Cross-functional partnership: Collaborate with Product, Engineering, Data, and Security to deliver external and internal value.
- Harden reliability: Improve observability (logging, metrics, tracing), and automated remediation to increase availability and latency.
- Automate everything: Use infrastructure-as-code and configuration management to make systems and processes repeatable, auditable, and secure.
- Scale cost-effectively: Optimize cluster utilization, autoscaling, and storage/networking to balance performance, reliability, and spend.
- Level-up developer experience: Build internal tooling, templates, and golden paths that reduce cognitive load and time-to-first-deploy for product teams.
- On-call & incident response: Participate in a sustainable on-call rotation, drive post-mortems, eliminate toil, and reduce MTTR via automation.
- Enable fast, safe delivery: Evolve CI/CD pipelines (build/test/release), and environment strategies (dev/stage/prod).
- AI: Build using agentic tools (Claude Code, Codex, etc) and push the boundaries of agentic development.
Requirements
- 5+ years of experience in software engineering with a focus on infrastructure, DevOps, and/or platform engineering.
- Team focused mindset, with solid collaboration and communication skills, with a focus on enabling others.
- Pragmatic problem-solver who communicates clearly, documents well, and thrives in fast-moving, high-ownership environments.
- Experience working with cloud infrastructure, specifically Kubernetes.
- Understanding of observability: metrics, logs, traces, and building actionable alerts/SLOs.
- Familiarity with infrastructure-as-code tools.
- Some programming experience in at least one modern programming language.
- Awareness of security fundamentals: IAM, workload identity, network policies, encryption, and secrets management.
Nice to Haves
- Open source contributions.
- Experience with company transitioning from startup to high-growth.
- Google Cloud Platform.
- Terraform.
- Python, Go, and/or JavaScript (TypeScript).
- Building and managing CI/CD systems and developer tooling.
Skills
Kubernetes, Terraform, GCP, Python, Go, JavaScript, TypeScript, CI/CD, Infrastructure As Code, Observability
Similar jobs
DevOps / SRE jobsOwn the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.