# Senior Software Engineer, Machine Learning Platform

**Company:** [Chime](https://hotfix.jobs/companies/chime)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $187k – $259k/yr
**Experience:** 5+ years
**Skills:** AWS, Machine Learning, Distributed Systems, Ray, Spark, Terraform, Kinesis, Kafka, Apache Flink, CI/CD, Docker, Kubernetes, Python, Go, Scala
**Posted:** 2026-09-11

> Build and operate scalable ML and AI platform infrastructure for training, inference, LLM, and agentic workloads on AWS. The role requires 5+ years of experience in ML infrastructure, platform engineering, distributed systems, or production ML systems, plus strong programming and DevOps skills.

## Job Description

## Responsibilities
- Design, build, and operate scalable ML and AI infrastructure on AWS.
- Build shared platform capabilities for LLM and agentic workloads, including model access, prompt and configuration lifecycle, retrieval, tool integration, state management, and workflow orchestration.
- Build evaluation frameworks for non-deterministic AI systems, including offline benchmarks, regression testing, online quality signals, human feedback, and failure analysis.
- Establish observability, reliability, and governance for models and agents, covering traces, model and prompt versions, tool calls, latency, token usage, quality, safety, privacy, and cost.
- Guide architecture decisions across traditional ML, LLM-powered applications, and agentic workflows, and contribute to the technical roadmap.
- Build distributed training, batch inference, and large-scale processing systems using Ray or Spark.
- Build and maintain infrastructure as code using Terraform.
- Support and evolve feature stores and feature pipelines.
- Develop data ingestion and streaming systems using Kinesis, Kafka, Flink, or Spark.
- Improve CI/CD workflows for ML models, AI applications, and platform components.
- Partner with Data Science and ML Engineering teams to improve developer experience.
- Participate in on-call rotations for production systems.

## Requirements
- Knowledge of the machine learning development lifecycle, including data preprocessing, model training, evaluation, deployment, and monitoring.
- Experience designing distributed systems and large-scale data or compute platforms on AWS using Spark or Ray.
- **5+ years of experience** in ML or AI infrastructure, platform engineering, distributed systems, or production ML systems.
- Working knowledge of LLM application patterns such as retrieval-augmented generation, structured outputs, tool calling, agent orchestration, and evaluation of non-deterministic systems.
- Experience designing production systems that integrate ML or foundation models through reliable APIs, workflows, and data contracts.
- Hands-on experience with CI/CD pipelines, DevOps practices, and infrastructure as code.
- Experience with containerization and orchestration technologies such as Docker and Kubernetes.
- Strong programming skills in Python, Go, Scala, Java, or similar languages.
- Understanding of software engineering fundamentals, including testing, version control, code review, and observability.

## Nice-to-haves
- Experience shipping LLM-powered or agentic systems to production.
- Experience with model gateways, prompt lifecycle management, retrieval or vector search, tool execution, or agent orchestration frameworks.
- Experience building evaluation, tracing, and observability capabilities for non-deterministic AI systems.
- Familiarity with managed or self-hosted foundation model infrastructure, such as Amazon Bedrock or SageMaker.
- Experience operating GPU-based workloads and optimizing training or inference performance and cost; CUDA experience is a plus.

## Compensation and Benefits
- Base salary: **$187,000–$259,000 USD annually**.
- Full-time employees are eligible for a bonus, equity package, and benefits.
- Benefits include health, financial, and wellbeing benefits; paid time off; parental leave; family planning reimbursement; commuter benefits; backup care; and an annual wellness stipend.

## Similar jobs

- [Senior Applied AI Engineer](https://hotfix.jobs/jobs/a0e924dc-c4ec-4b69-ad92-52521d004065) - Mintlify - San Francisco, CA - $190k – $265k/yr
- [Senior AI Engineer - FDE - U.S. Federal Sector](https://hotfix.jobs/jobs/f4e6bcfe-2e0e-4d0b-b21e-757eb2df2dd6) - Databricks - Remote - $182k – $250k/yr
- [Senior ML/AI Modeler, Risk Automation Machine Learning](https://hotfix.jobs/jobs/303b4f2e-9bc8-4034-8fa1-d17c71bf6f1e) - Square - Remote - $195k – $343k/yr
- [Senior Software Engineer, Autonomy Capabilities](https://hotfix.jobs/jobs/ae9ec2ef-285e-42da-9740-0b3ce5faf8cf) - Shield AI - San Mateo, CA - $196k – $294k/yr
- [Senior Software Engineer: AI](https://hotfix.jobs/jobs/99592ad8-7306-4ebe-b328-2713eee97510) - Swayable - New York, NY - $175k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/f3baee7b-9c39-412b-adde-c3a6a76fffde
**Canonical:** https://hotfix.jobs/jobs/f3baee7b-9c39-412b-adde-c3a6a76fffde