Skip to content
ChimeChime

Senior Software Engineer, Machine Learning Platform

Build and operate scalable ML and AI platform infrastructure for training, inference, LLM, and agentic workloads on AWS. The role requires 5+ years of experience in ML infrastructure, platform engineering, distributed systems, or production ML systems, plus strong programming and DevOps skills.

About the job

Responsibilities

  • Design, build, and operate scalable ML and AI infrastructure on AWS.
  • Build shared platform capabilities for LLM and agentic workloads, including model access, prompt and configuration lifecycle, retrieval, tool integration, state management, and workflow orchestration.
  • Build evaluation frameworks for non-deterministic AI systems, including offline benchmarks, regression testing, online quality signals, human feedback, and failure analysis.
  • Establish observability, reliability, and governance for models and agents, covering traces, model and prompt versions, tool calls, latency, token usage, quality, safety, privacy, and cost.
  • Guide architecture decisions across traditional ML, LLM-powered applications, and agentic workflows, and contribute to the technical roadmap.
  • Build distributed training, batch inference, and large-scale processing systems using Ray or Spark.
  • Build and maintain infrastructure as code using Terraform.
  • Support and evolve feature stores and feature pipelines.
  • Develop data ingestion and streaming systems using Kinesis, Kafka, Flink, or Spark.
  • Improve CI/CD workflows for ML models, AI applications, and platform components.
  • Partner with Data Science and ML Engineering teams to improve developer experience.
  • Participate in on-call rotations for production systems.

Requirements

  • Knowledge of the machine learning development lifecycle, including data preprocessing, model training, evaluation, deployment, and monitoring.
  • Experience designing distributed systems and large-scale data or compute platforms on AWS using Spark or Ray.
  • 5+ years of experience in ML or AI infrastructure, platform engineering, distributed systems, or production ML systems.
  • Working knowledge of LLM application patterns such as retrieval-augmented generation, structured outputs, tool calling, agent orchestration, and evaluation of non-deterministic systems.
  • Experience designing production systems that integrate ML or foundation models through reliable APIs, workflows, and data contracts.
  • Hands-on experience with CI/CD pipelines, DevOps practices, and infrastructure as code.
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes.
  • Strong programming skills in Python, Go, Scala, Java, or similar languages.
  • Understanding of software engineering fundamentals, including testing, version control, code review, and observability.

Nice-to-haves

  • Experience shipping LLM-powered or agentic systems to production.
  • Experience with model gateways, prompt lifecycle management, retrieval or vector search, tool execution, or agent orchestration frameworks.
  • Experience building evaluation, tracing, and observability capabilities for non-deterministic AI systems.
  • Familiarity with managed or self-hosted foundation model infrastructure, such as Amazon Bedrock or SageMaker.
  • Experience operating GPU-based workloads and optimizing training or inference performance and cost; CUDA experience is a plus.

Compensation and Benefits

  • Base salary: $187,000–$259,000 USD annually.
  • Full-time employees are eligible for a bonus, equity package, and benefits.
  • Benefits include health, financial, and wellbeing benefits; paid time off; parental leave; family planning reimbursement; commuter benefits; backup care; and an annual wellness stipend.

Skills

AWS, Machine Learning, Distributed Systems, Ray, Spark, Terraform, Kinesis, Kafka, Apache Flink, CI/CD, Docker, Kubernetes, Python, Go, Scala

Mintlify

Mintlify

San Francisco, CA

Senior Applied AI Engineer
$190k+/yrOn-site4+ YOEML Engineering

Senior engineer building and leading production AI products, including agent systems, evaluations, and full-stack customer experiences. Requires 4+ years of software development experience, deep language-model production expertise, and ownership from experimentation through deployment.

Databricks

Databricks

United States

Senior AI Engineer - FDE - U.S. Federal Sector
$182k+/yrRemote7+ YOEML Engineering

Build and productionize generative AI applications for U.S. federal customers, advise clients, and influence product direction. The role requires extensive data science and machine learning deployment experience, a graduate quantitative degree or equivalent experience, and U.S. security clearance eligibility.

Square

Square

United States

Senior ML/AI Modeler, Risk Automation Machine Learning
$195k+/yrRemote8+ YOEML Engineering

Leads the design, deployment, and optimization of agentic and generative AI systems that automate risk and compliance investigations at scale. Requires 8+ years of machine learning modeling experience, production ML expertise, and advanced technical education.

Shield AI

Shield AI

San Mateo, CA

Senior Software Engineer, Autonomy Capabilities
$196k+/yrOn-site5+ YOEML Engineering

Design and deploy tactical autonomy algorithms and high-performance software for unmanned systems operating in complex, contested environments. The role requires 5+ years of related experience, strong C++ and Python skills, robotics expertise, and the ability to obtain a SECRET clearance.

Swayable

Swayable

New York, NY
Senior Software Engineer: AI
$175k+/yrOn-site5+ YOEML Engineering

Senior engineer developing and productizing AI, machine learning, scientific computing, and data-analysis capabilities for a high-performance analytics engine. Requires 5+ years building quantitative data-intensive software and expertise in Python, machine learning, scalable architecture, and distributed computing.