Senior Platform Engineer
Build and operate AI-assisted platform infrastructure, developer tooling, and production workflows across Kubernetes, AWS, GitOps, observability, and FinOps. Requires 4+ years of platform, infrastructure, or backend engineering experience and hands-on experience with MCP servers and agentic AI systems.
About the job
Responsibilities
- Design and build MCP servers that expose platform capabilities as safe, well-scoped tools for AI agents and developer-facing assistants.
- Contribute to enterprise MCP patterns, LLM tooling, agentic guardrails, knowledge repositories, and framework rollout.
- Design structured, agentic workflows for incident triage, deployment validation, configuration remediation, and capacity planning.
- Drive AI-assisted code review and developer tooling using tool-use, validation steps, and human-in-the-loop gates.
- Operationalize LLM-based features with structured prompting, RAG, output validation, and evaluation harnesses.
- Implement LLM gateway and router patterns to control costs, route models, and provide observability for AI workloads.
- Design and evolve GitOps-based continuous delivery and Kubernetes infrastructure on AWS EKS.
- Manage infrastructure as code with Terraform, Helm, and Kustomize.
- Strengthen observability, reliability, and operational excellence through SLOs, error budgets, metrics, traces, logs, and automation that improves MTTD and MTTR.
- Extend the developer control plane with paved paths, scorecards, and self-service actions.
- Instrument platform cost signals, including compute, observability, and AI/LLM spend, and build FinOps automation.
- Define success metrics and run time-bound experiments evaluating developer efficiency, reliability, and cost.
- Document usage guidance, patterns, and best practices for proven AI workflows.
Requirements
- Bachelor's or master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 4+ years of platform, infrastructure, or backend engineering experience operating production systems in a cloud environment; AWS preferred.
- Strong Python and/or Go coding skills with experience building and operating production services.
- Deep experience with Kubernetes, preferably EKS; GitOps with Argo CD; CI/CD with GitHub Actions; and Terraform, Helm, and Kustomize.
- Experience with service mesh technologies.
- Strong observability skills with Datadog APM, metrics, tracing, and logs, plus a record of improving reliability and promoting SLO/error-budget practices.
- Hands-on experience building or integrating MCP servers, including tool-surface design, scope and authorization management, and connecting AI agents to infrastructure.
- Production experience with structured or agentic AI workflows, including planning/execution separation, human-in-the-loop validation, RAG, and tool-use patterns.
- FinOps experience covering cloud cost attribution, workload optimization, and AI/LLM spend control.
- Familiarity with LLM gateway and router patterns for cost control, model routing, and AI workload observability.
- Clear communication and effective collaboration with distributed partner teams.
- Experience using AI-assisted development tools such as GitHub Copilot, Cursor, or ChatGPT.
Benefits
- Healthcare coverage.
- Internet and cell phone reimbursement.
- Learning and development stipend.
- Potential opportunities to travel to the Mountain View headquarters.
- Hybrid work based in the Bengaluru office, with two days on-site.
- Visa sponsorship and immigration support are not available.
Skills
Python, Go, Kubernetes, Aws Eks, Argo Cd, GitHub Actions, Terraform, Helm, Kustomize, Datadog, Mcp Servers, RAG, Llm Gateways, Finops, GitOps
Similar jobs
ML Engineering jobsBuild and ship autonomous, agentic software development lifecycle capabilities, including AI agents, orchestration, and safety guardrails. The role requires senior software engineering experience, proficiency in Ruby, Go, or Python, distributed systems knowledge, and experience with AI/ML applications.
Build and operate scalable AI/ML systems and pipelines that strengthen Airbnb’s fraud prevention and trust defenses. The role requires 7+ years of backend or platform engineering experience, strong programming and data engineering skills, and production machine learning expertise.
Build production AI agents and the distributed platform that operates large-scale GPU infrastructure. The role requires 5+ years of backend, distributed systems, or infrastructure experience, with expertise in agent systems, knowledge graphs, retrieval, or semantic search.
Build and deploy explainable machine learning, NLP, LLM, and agentic systems that power enterprise go-to-market intelligence products. The role requires 6+ years of production ML experience, strong Python and cloud skills, and end-to-end ownership from modeling through monitoring.
Senior software engineer developing ML-based search relevance and discovery systems, including query understanding, ranking, retrieval, and evaluation pipelines. The role requires 5+ years of search relevance experience and expertise in NLP, LLMs, or related discovery technologies.