Skip to content
LlamaIndexLlamaIndex

Senior Infra Engineer

Builds and scales core infrastructure including Kubernetes clusters and cloud resources to power AI data platform. Collaborates with teams on foundational systems, optimizes for cost/performance, and ensures security/compliance. Requires 5+ years infra experience.

About the job

Responsibilities

  • Collaborate with other engineering teams to build and maintain foundational systems that empower developers and support the company's rapid growth.
  • Design and implement scalable infrastructure solutions for various deployment models, including SaaS, single-tenant, and private deployments.
  • Manage and optimize cloud resources and Kubernetes clusters for cost-effectiveness and performance.
  • Enable external customer deployment success through maintaining clear infrastructure boundaries and principles.
  • Optimize and improve the release and deployment processes to enhance efficiency and reliability.
  • Ensure compliance with relevant regulations and implement robust security measures across different deployment environments.

Qualifications

  • 5+ years of engineering experience.
  • Worked on Platform or Infrastructure teams on significant projects involving infrastructure components (Terraform/CDKTF, Kubernetes, Helm, test infrastructure, release management, observability, etc.).
  • Experience in optimizing cloud resource utilization. Proficient in tuning Kubernetes clusters and cloud resources for cost and performance efficiency.
  • Willing to build LlamaIndex's engineering culture as we grow.
  • You can balance speed and pragmatism and build the appropriate solutions for each stage of the company's growth.

Preferred Qualifications

  • Experience building out infrastructure from the ground up at a fast-growing startup.
  • Experience with observability tools like Prometheus, Grafana, New Relic.
  • Experience with GitOps tools like ArgoCD and Flux for continuous deployment.
  • Experience with security compliance and audits in cloud environments such as SOC2.
  • Familiar with Python, Postgres, multi-cloud deployments.

Skills

Kubernetes, Terraform, Helm, Prometheus, Grafana, Argo CD, Flux, Python, Postgres, Cdktf

Runpod

Runpod

United States

Senior HPC Storage Engineer
$180k+/yrRemote8+ YOEDevOps / SRE

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.

Onebrief

Onebrief

Colorado Springs, CO

Senior Site Reliability Engineer, Colorado Springs
$180k+/yrOn-site5+ YOEDevOps / SRE

Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.