Staff Software Engineer – AI Platform
Build and operate backend services, distributed systems, and cloud infrastructure for AI Platform products. The role requires 8+ years of software engineering experience, strong production backend development, Terraform and Kubernetes expertise, and the ability to mentor engineers and drive platform reliability.
About the job
Responsibilities
- Design, build, and operate backend services and infrastructure powering AI Platform products.
- Write maintainable, reliable, and performant production code for platform and service layers.
- Architect and evolve distributed systems supporting real-time interactions, asynchronous processing, and secure data workflows.
- Build and extend cloud-native infrastructure across core services, networking, storage, identity, and observability layers.
- Design and maintain infrastructure as code using Terraform, including reusable modules, environment management, and drift-safe workflows.
- Build and operate Kubernetes/EKS-based services, including deployment strategies, health checks, scaling configuration, and containerization workflows.
- Design and improve CI/CD pipelines, release workflows, and engineering productivity systems.
- Build and maintain asynchronous and event-driven systems using queues, RPC/service communication, and resilient failure-handling patterns.
- Design and evolve relational and NoSQL data models, including schema design, migrations, replication, and access patterns.
- Build and maintain observability across metrics, logs, tracing, dashboards, and alerting.
- Apply secure-by-default practices across IAM, secrets, encryption, network isolation, and production operations.
- Mentor engineers through code reviews, design reviews, incident reviews, and technical guidance.
- Drive technical decisions that improve platform reliability, scalability, developer productivity, and long-term system health.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- At least 8+ years of professional software engineering experience, including significant backend, infrastructure, or platform-oriented work.
- Strong proficiency in at least one production backend language such as Python, Go, or Java, with the ability to work across a polyglot environment.
- Strong software engineering fundamentals, including system design, debugging, code quality, testing, and operational ownership.
- Deep hands-on experience with cloud infrastructure in production environments.
- Strong experience with infrastructure-as-code workflows such as Terraform or equivalent tools.
- Experience with Kubernetes/EKS, Docker, and modern deployment practices.
- Experience designing or extending CI/CD pipelines in production environments.
- Experience with distributed systems, microservice communication, message queues, asynchronous processing, and failure handling.
- Experience with relational and NoSQL database design in production systems.
- Strong observability and production-debugging skills.
- Strong communication, collaboration, and mentoring skills.
- Ability to operate as a software engineer first, with substantial production software engineering depth.
Nice to Have
- Experience with AI/ML platform tooling such as Databricks, Spark, MLflow, or model-serving systems.
- Experience with real-time streaming or WebSocket-based systems at scale.
- Experience with multi-region architectures, data residency, or failover design.
- Experience in fintech or other highly regulated environments.
- Experience optimizing platform cost, performance, and reliability tradeoffs at scale.
Work Arrangement
- Hybrid role requiring work from the Pune office 3 days per week.
Skills
Python, Go, Java, Terraform, Kubernetes, Amazon Eks, Docker, CI/CD, Distributed Systems, Message Queues, Relational Databases, Nosql Databases, Observability, IAM, MLflow
Similar jobs
Backend Engineering jobsLeads architecture and reliability for a global payments and commerce platform, shaping APIs, distributed systems, observability, compliance, and production operations. Requires 10+ years of backend experience, deep Go and distributed-systems expertise, and Staff/Principal-level technical leadership.
Leads architecture and hands-on development of scalable, reliable, and compliant financial integrity systems. The role requires 8+ years of software engineering experience, distributed-systems expertise, and the ability to drive complex initiatives and mentor engineers through influence.
Provides cross-team technical leadership for large backend initiatives, modular architecture, production operations, and AI-assisted engineering adoption. The role requires deep backend architecture experience, monolith modernization expertise, strong written influence, and mentoring ability.
Staff Software Engineer responsible for architecting and scaling security platform services, defining technical strategy, and mentoring engineers. Requires at least 8 years of software engineering experience plus strong system design, coding, and production-service expertise.
Staff Engineer leading design and delivery of large-scale distributed data platform features for enterprise customers. The role requires strong Java and cloud experience, architectural leadership, cross-functional collaboration, and mentorship.