Principal Software Engineer — Backend & Infrastructure
Leads backend and ML infrastructure architecture across distributed processing, GPU fleets, model serving, and reliability. The role requires 10+ years of large-scale systems experience, strong architectural ownership, and a Computer Science degree or equivalent track record.
About the job
Responsibilities
- Own the architecture for real-time data processing at scale, including distributed messaging systems for high-throughput streaming workloads with strict latency requirements.
- Build and guide the ML platform roadmap for training and serving infrastructure.
- Scale GPU infrastructure through capacity planning, scheduling, utilization optimization, and burst-demand management across training and inference fleets.
- Improve inference latency and cost per request through batching, routing, quantization, compilation, autoscaling, and caching while maintaining output quality.
- Drive reliability, uptime, observability, and incident response for production serving systems.
- Turn ambiguous business problems into executable technical plans with Product and GTM partners.
- Lead design reviews, mentor senior engineers, and establish durable engineering practices.
- Evaluate emerging tools and techniques and adopt them selectively.
Requirements
- 10+ years building backend and infrastructure systems, with experience owning architecture and design at scale.
- Deep hands-on experience with large-scale databases, high-throughput messaging systems, and real-time job queues.
- Ability to navigate large, complex codebases and reason about architectural tradeoffs.
- Experience mentoring senior engineers and driving technical decisions through influence.
- Strong written communication skills for technical and executive audiences.
- BTech, MTech, PhD in Computer Science, or equivalent experience.
Nice-to-haves
- Production experience with Django, Celery, Redis, PostgreSQL, and Google Cloud.
- Experience scaling GPU infrastructure and model inference in production, including capacity planning, scheduling, utilization, autoscaling, and latency/cost optimization.
- Experience with vLLM, TensorRT, Triton, Ray Serve, or equivalent inference tooling.
- Experience scaling a platform through a comparable growth stage, particularly at a global product company's India site.
- Background in speech, NLP, or information retrieval systems.
Compensation & Benefits
- Health coverage for the employee, spouse, children, and parents.
- Real ownership over systems used by enterprises worldwide, with autonomy over their design and implementation.
Skills
Python, Django, Celery, Redis, Postgres, GCP, Distributed Systems, Messaging Systems, Real-Time Processing, Gpu Infrastructure, Machine Learning Infrastructure, Model Inference, vLLM, TensorRT, Kubernetes
Similar jobs
ML Engineering jobsLeads the development of ML- and NLP-powered search relevance systems, including query understanding, ranking, retrieval, and evaluation. Requires 10+ years of search relevance experience and a bachelor’s degree, with advanced study preferred.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Leads the development of machine-learning search relevance systems, including query understanding, ranking, retrieval, and evaluation at scale. Requires 10+ years of search relevance experience and expertise in ML, NLP, or related discovery technologies.
Leads the development of machine-learning search relevance systems, including query understanding, ranking, retrieval, and evaluation pipelines. The role requires 10+ years of search relevance experience and expertise in NLP, LLMs, or related discovery technologies.
Build and ship autonomous, agentic software development lifecycle capabilities, including AI agents, orchestration, and safety guardrails. The role requires senior software engineering experience, proficiency in Ruby, Go, or Python, distributed systems knowledge, and experience with AI/ML applications.