Senior Backend Engineer
Architects and builds high-performance Golang backend services for AI PaaS platform, focusing on distributed systems, GPU orchestration, model deployment, and cloud-native infrastructure. Requires 5+ years experience with expert Golang and Kubernetes proficiency.
About the job
What You'll Do
Core Platform & Infrastructure Backend
- Architect and develop high-performance Golang services for FlexAI's AI PaaS and infrastructure platform
- Build internal APIs powering model deployment, job scheduling, and compute lifecycle management
- Develop components interfacing with GPU/compute infrastructure and AI runtimes
Distributed Systems & Scalability
- Design and scale microservices and event-driven systems for high-throughput AI workloads
- Optimize for low latency, high concurrency, and fault tolerance
- Implement service-to-service communication (gRPC/REST, message queues, async pipelines)
- Drive reliability, observability, and resilience across services
AI Platform Integration
- Collaborate with AI/ML and Runtime teams to integrate systems with training pipelines, inference infrastructure, experimentation workflows, and dataset/artifact management
- Enable orchestration across cloud and on-prem environments
- Build abstractions that simplify AI infrastructure consumption
Cloud-Native & Platform Engineering
- Design cloud-native, Kubernetes-native services
- Work with DevOps/SRE on CI/CD, deployment automation, and scalability
- Contribute to architecture decisions for multi-region, multi-cloud infrastructure
- Improve monitoring, logging, and diagnostics
Technical Leadership
- Lead architecture reviews and set engineering standards
- Mentor engineers and guide complex problem-solving
- Drive long-term roadmap for backend infrastructure and AI platform capabilities
- Partner with Product, Runtime, and Infra leadership to translate requirements into scalable systems
Tech Stack (Indicative):
Languages: Golang (Primary), Python (Secondary)
Infrastructure: Kubernetes, Docker, Cloud (AWS/GCP/Azure)
Architecture: Microservices, gRPC, Event-driven systems
Data: SQL + NoSQL databases, caching, streaming systems
Observability: Prometheus, Grafana, OpenTelemetry (or similar)
What You'll Need to Be Successful
Core Engineering
- 5+ years of Backend or Infrastructure Engineering experience
- Expert-level proficiency in Golang (must-have, heavy hands-on)
- Strong experience building production-grade distributed systems
- Proven track record on infrastructure platforms, PaaS, or deep-tech systems
Infrastructure & Systems
- Deep understanding of cloud-native architectures and containerized environments
- Strong experience with Kubernetes, Docker, and cluster orchestration
- Familiarity with compute scheduling, resource management, or platform runtimes is a strong plus
Databases & Data Systems
- Experience with distributed databases (PostgreSQL, Cassandra, DynamoDB, etc.)
- Strong understanding of caching, queues, and streaming systems (Redis, Kafka, etc.)
AI / Platform Exposure (Highly Preferred)
- Experience on AI/ML platforms, model infrastructure, or data platforms
- Familiarity with ML pipelines, inference systems, or GPU-backed workloads
- Exposure to PyTorch, TensorFlow infrastructure, or model serving systems is a plus
Skills
Go, Kubernetes, Docker, gRPC, Postgres, Cassandra, DynamoDB, Redis, Kafka, Prometheus, Grafana, OpenTelemetry, AWS, GCP, Azure
Similar jobs
Backend Engineering jobsBuild and operate high-throughput blockchain infrastructure, APIs, and platform primitives integrating protocols such as Ethereum and Bitcoin with internal services. Requires 5+ years of software engineering experience, distributed-systems expertise, and hands-on crypto infrastructure experience.
Design, build, and operate Cloudflare’s globally distributed cache and reverse-proxy data plane, improving performance, correctness, and resilience across the edge. Requires at least 4 years of production systems experience and proficiency in a systems or backend language.
Senior backend engineer designing and operating reliable billing and financial systems, APIs, data models, and distributed workflows. The role requires 5+ years of professional software development experience, strong backend expertise, and collaboration across Product, Finance, Operations, and Data.
Senior individual contributor responsible for designing, building, operating, and improving large-scale backend services, APIs, and telemetry pipelines in Go and Python. The role requires production systems ownership, distributed-systems expertise, incident leadership, mentoring, and technical design leadership.
Build and operate backend services, data pipelines, storage, and retrieval systems that provide trusted context to agentic platforms and product applications. The role requires 8+ years of software engineering experience, distributed-systems expertise, cloud infrastructure knowledge, and strong data modeling skills.