Senior Infrastructure Engineer
Build and operate multi-tenant, distributed infrastructure for an enterprise AI platform across Kubernetes and multiple clouds. The role requires 5+ years of platform, infrastructure, or backend engineering experience, strong Golang and Python skills, and deep expertise in reliability, databases, and distributed systems.
About the job
Responsibilities
- Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation with namespaces, network policies, per-tenant resource quotas, and RBAC.
- Build core platform and data-plane components in Golang and Python, including data ingestion, knowledge-base indexing, vector and graph search, application connectivity, workflow automation, and ML operations, against explicit latency and throughput SLOs.
- Own service-to-service communication, including gRPC/Protobuf API contracts, service meshes such as Istio and Linkerd, load balancing, retries, timeouts, and circuit breaking.
- Make and document architectural tradeoffs involving partitioning and sharding, consistency models, caching tiers, and build-versus-buy decisions.
- Define reliability contracts including SLIs, SLOs, error budgets, capacity planning, autoscaling, and graceful degradation.
- Design and operate observability using Prometheus, Grafana, OpenTelemetry, distributed tracing, and real-time alerting.
- Drive DevOps and platform-engineering practices using infrastructure as code, Helm, GitOps, and CI/CD pipelines.
- Optimize performance and cost through profiling, load testing, latency budgets, and cost-per-request analysis.
- Participate in on-call rotations and lead incident response and root-cause analysis.
Requirements
- Bachelor's degree in Computer Science or a related field.
- 5+ years of experience in platform, infrastructure, or backend engineering.
- Strong computer science fundamentals, including data structures, algorithms, operating systems, and networking.
- Proficiency in Golang and Python.
- Production experience with Docker, Kubernetes, and microservices architecture.
- Hands-on experience with GCP, Azure, or AWS; multi-cloud experience is a strong plus.
- Strong database expertise, including query and read/write-path optimization, partitioning and sharding, replication, consistency models, NoSQL, graph stores, the CAP theorem, and database internals.
- Distributed-systems experience with idempotency, backpressure, delivery semantics, and message queues such as Kafka, Pulsar, NATS, or Pub/Sub.
- Track record of building platforms from the ground up for use by other engineering teams.
Nice-to-haves
- Experience operating systems at high scale, including high QPS and large data volumes.
- Experience with secrets management, mTLS, RBAC, OIDC/SAML, and network policy.
- Experience with vector databases such as pgvector, Pinecone, or Milvus, and graph databases such as Neo4j or Neptune.
- Open-source contributions to infrastructure projects, such as Kubernetes operators.
Compensation and Benefits
- Compensation is determined by location, level, job-related knowledge, skills, and experience.
- Certain roles may be eligible for variable compensation, equity, and benefits.
Skills
Kubernetes, Go, Python, Docker, GCP, Azure, AWS, Microservices, gRPC, Istio, Linkerd, Terraform, Helm, Argo CD, Prometheus
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.
Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.
Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.