Skip to content
EmaEma

Platform Engineer

Builds and operates multi-tenant, cloud-native platform infrastructure for an agentic AI product. The role requires 5+ years of platform, infrastructure, or backend engineering experience, strong Golang and Python skills, Kubernetes expertise, and deep distributed-systems knowledge.

About the job

Responsibilities

  • Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation with namespaces, network policies, per-tenant resource quotas, and RBAC.
  • Build core platform and data-plane components in Golang and Python, including data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and MLOps, against explicit latency and throughput SLOs.
  • Own service-to-service communication, including gRPC/protobuf API contracts, service meshes, load balancing, retries, timeouts, and circuit breaking.
  • Document architectural tradeoffs involving partitioning and sharding, consistency models, caching tiers, and build-versus-buy decisions.
  • Define reliability contracts covering SLIs/SLOs, error budgets, capacity planning, autoscaling, and graceful degradation.
  • Design and operate observability systems for system health visibility.
  • Drive DevOps and platform-engineering practices, including infrastructure as code, Helm, GitOps, and CI/CD pipelines.
  • Optimize performance and cost through profiling, load testing, latency budgets, and cost-per-request analysis.
  • Participate in on-call rotations and lead incident response and root-cause analysis.

Requirements

  • Bachelor's degree in Computer Science or a related field.
  • 5+ years of experience in Platform, Infrastructure, or Backend Engineering.
  • Strong computer science fundamentals, including data structures, algorithms, operating systems, and networking.
  • Proficiency in Golang and Python.
  • Production experience with Docker, Kubernetes, and microservices architecture.
  • Hands-on experience with GCP, Azure, or AWS; multi-cloud experience is a strong plus.
  • Strong database expertise, including query and read/write-path optimization, partitioning and sharding, replication, consistency models, NoSQL, and graph stores.
  • Understanding of the CAP theorem and database internals.
  • Distributed-systems experience with idempotency, backpressure, delivery semantics, and message queues.
  • Track record of building platforms that other engineering teams successfully build on.

Nice-to-haves

  • Experience operating systems at high scale, including high QPS and large data volumes.
  • Experience with authentication and security, including Vault, mTLS, RBAC, OIDC/SAML, and network policies.
  • Experience with vector databases such as pgvector, Pinecone, or Milvus.
  • Experience with graph databases such as Neo4j or Neptune.
  • Open-source contributions to infrastructure projects, such as Kubernetes operators.

Compensation and Benefits

  • Compensation is determined by location, level, job-related knowledge, skills, and experience.
  • Certain roles may be eligible for variable compensation, equity, and benefits.

Skills

Go, Python, Kubernetes, Docker, GCP, Azure, AWS, Microservices, gRPC, Istio, Linkerd, Terraform, Helm, Argo CD, Prometheus

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Acryldata

Acryldata

Bengaluru, India

DevOps
No salary listedRemote5+ YOEDevOps / SRE

Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.

Invisible Tech

Invisible Tech

Estonia
Site Reliability Engineer
No salary listedRemoteDevOps / SRE

Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.