# Principal Software Engineer — Backend & Infrastructure

**Company:** [Level AI](https://hotfix.jobs/companies/level-ai)
**Location:** Noida, India, Bengaluru, India
**Role:** ML Engineering
**Experience:** 10+ years
**Skills:** Python, Django, Celery, Redis, Postgres, GCP, Distributed Systems, Messaging Systems, Real-Time Processing, Gpu Infrastructure, Machine Learning Infrastructure, Model Inference, vLLM, TensorRT, Kubernetes
**Posted:** 2026-08-10

> Leads backend and ML infrastructure architecture across distributed processing, GPU fleets, model serving, and reliability. The role requires 10+ years of large-scale systems experience, strong architectural ownership, and a Computer Science degree or equivalent track record.

## Job Description

## Responsibilities
- Own the architecture for real-time data processing at scale, including distributed messaging systems for high-throughput streaming workloads with strict latency requirements.
- Build and guide the ML platform roadmap for training and serving infrastructure.
- Scale GPU infrastructure through capacity planning, scheduling, utilization optimization, and burst-demand management across training and inference fleets.
- Improve inference latency and cost per request through batching, routing, quantization, compilation, autoscaling, and caching while maintaining output quality.
- Drive reliability, uptime, observability, and incident response for production serving systems.
- Turn ambiguous business problems into executable technical plans with Product and GTM partners.
- Lead design reviews, mentor senior engineers, and establish durable engineering practices.
- Evaluate emerging tools and techniques and adopt them selectively.

## Requirements
- 10+ years building backend and infrastructure systems, with experience owning architecture and design at scale.
- Deep hands-on experience with large-scale databases, high-throughput messaging systems, and real-time job queues.
- Ability to navigate large, complex codebases and reason about architectural tradeoffs.
- Experience mentoring senior engineers and driving technical decisions through influence.
- Strong written communication skills for technical and executive audiences.
- BTech, MTech, PhD in Computer Science, or equivalent experience.

## Nice-to-haves
- Production experience with Django, Celery, Redis, PostgreSQL, and Google Cloud.
- Experience scaling GPU infrastructure and model inference in production, including capacity planning, scheduling, utilization, autoscaling, and latency/cost optimization.
- Experience with vLLM, TensorRT, Triton, Ray Serve, or equivalent inference tooling.
- Experience scaling a platform through a comparable growth stage, particularly at a global product company's India site.
- Background in speech, NLP, or information retrieval systems.

## Compensation & Benefits
- Health coverage for the employee, spouse, children, and parents.
- Real ownership over systems used by enterprises worldwide, with autonomy over their design and implementation.

## Similar jobs

- [Staff Software Engineer - Search Platform](https://hotfix.jobs/jobs/9c9d7c76-3610-4456-81fe-2662b6af213e) - Databricks - Bengaluru, India
- [Staff Machine Learning Engineer](https://hotfix.jobs/jobs/2cb6da8f-4a54-448d-a4b6-fc05f1412a12) - Payabli - Remote
- [Staff Software Engineer](https://hotfix.jobs/jobs/4f600fb6-7f6e-4718-b395-16f02be6a016) - Databricks - Bengaluru, India
- [Staff Software Engineer - Machine Learning](https://hotfix.jobs/jobs/b2e23245-bfe9-48d4-9338-2b865f5b19f6) - Databricks - Bengaluru, India
- [Senior Software Engineer, NLP](https://hotfix.jobs/jobs/be3f0222-b08a-4620-9ddc-402209ea73f7) - GitLab - Remote

**Apply:** https://hotfix.jobs/jobs/510e8cbc-9aba-4d00-bd93-c981cb2a8812
**Canonical:** https://hotfix.jobs/jobs/510e8cbc-9aba-4d00-bd93-c981cb2a8812