Software Engineer - Platforms & Productivity
Build and operate internal developer platforms that improve engineering velocity, reliability, and security. The role spans developer tooling, CI/CD, GitOps, AI-assisted development, and automated engineering guardrails.
About the job
Responsibilities
- Build developer productivity tools, MCP servers, and AI agents.
- Automate Engineering Codex controls through CI/CD checks, templates, and policy-as-code.
- Monitor Codex adoption, improve controls, and support remediation and exception workflows.
- Improve GitOps, developer experience, platform security, and reliability.
- Partner with product teams on architecture and platform adoption.
- Respond to and prevent incidents affecting internal platforms.
Requirements
- Experience with TypeScript, Go, or Bash.
- Strong debugging and source-control skills.
- Experience with CI/CD, GitOps, or developer platforms.
- Ability to break down complex problems and evaluate trade-offs.
- Clear communication and a strong focus on developer experience.
Nice-to-haves
- Experience with AI agents, MCP, evals, or LLM-based tooling.
- Experience with policy-as-code or compliance automation.
- Background in platform engineering or developer productivity.
- Experience operating distributed or Kubernetes-based systems.
Skills
TypeScript, Go, Bash, CI/CD, GitOps, Mcp, AI Agents, Policy-As-Code, Kubernetes, Source Control
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.
Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.