Build, deploy, and operate scalable ML systems for real-time applications like anomaly detection, recommendations, and agentic AI at Twilio. Requires 5+ years production ML experience, strong Python/Java/SQL skills, and MLOps expertise.
156k – 229k/yr
Remote5+ YOEML Engineering
About the role
Responsibilities
Partner with product, UX, and technical stakeholders to analyze business problems, clarify requirements, define scope, and translate them into measurable ML problem statements.
Design, implement, and maintain scalable, enterprise-grade ML solutions in production.
Build reproducible ML workflows for data preparation, training, evaluation, and inference using modern orchestration and MLOps tooling.
Implement monitoring and evaluation frameworks to continuously improve data quality, model performance, latency, and cost through feedback loops.
Partner cross-functionally with Product, Data Science/ML, Engineering, and Security to deliver resilient, scalable, and compliant ML-powered services.
Demonstrate end-to-end systems understanding and articulate the “why” behind model and system design choices.
Own operational excellence: SLAs, on-call, incident response, customer feedback triage, and blameless post-mortems.
Drive engineering excellence via AI-assisted SDLC, code reviews, automated testing, MLOps best practices, knowledge-sharing, and mentoring.
Actively adopt AI-assisted practices to improve implementation and collaboration efficiency.
Qualifications
Required:
Strong foundation in ML/AI (statistics, probability, optimization) with the ability to apply these concepts to real-world problems.
5+ years of experience building, deploying, and operating data and ML systems in production.
Proficient in Python, Java, and SQL; strong software engineering fundamentals (system design, testing, version control, code reviews).
Hands-on experience with workflow orchestration and data pipelines (e.g., Airflow, Kubeflow) and cloud data platforms/storage (e.g., SageMaker Feature Store, Snowflake, DynamoDB, OpenSearch).
Experience with the ML lifecycle and MLOps tooling (e.g., MLflow, Metaflow, SageMaker; LLM/agent frameworks such as LangChain/LangGraph; model evaluation/observability tools such as Galileo or similar).
Working knowledge of containerization and cloud infrastructure, including Docker and Kubernetes, GitOps/CI/CD tools (e.g., Argo CD), and at least one major cloud platform (AWS, GCP, or Azure).
Understanding of data modeling and scalable systems, including distributed computing and streaming frameworks (e.g., Spark/EMR, Flink, Kafka Streams); familiarity with GPU-based implementation is a plus.
Demonstrated ability to ramp up quickly and operate effectively in new application/business domains.
Strong written and verbal communication skills: able to document and present designs and decisions, and comfortable giving/receiving feedback in an Agile environment.
Desired:
Familiarity with ML problem areas and techniques, including recommendation systems (e.g., graph-based approaches, two-tower models), time-series modeling (classical and deep learning), representation learning (e.g., embeddings), anomaly detection, and causal inference.
Practical experience with LLMs and generative AI workflows, including foundation model fine-tuning, RAG, and vector databases.
Evidence of technical leadership/impact, such as contributions to open-source data/ML projects and/or published technical presentations, blog posts, papers, or research.
Domain experience (plus) in communications, marketing automation, or customer engagement analytics.
Familiarity with AI-assisted development tools (e.g., Claude, GitHub Copilot/Codex, Cursor, etc.).
Advanced degree preferred (M.S. or Ph.D.) in a relevant field.
Build and improve AI agent experiences and conversational knowledge engine for Otter AI Chat. Focus on quality evaluation, infrastructure for orchestration/tracing, diagnosing failures across the stack, and driving measurable improvements in AI systems from traces and feedback. Requires 3+ years AI/ML engineering, strong backend/distributed systems skills, and experience shipping LLM-powered products.
155k – 185k/yr
Hybrid3+ YOEML Engineering
Applied AI Engineer
RampNew York, NY +1
Build and ship full-stack AI projects including AI agents, RAG, structured extraction, and LLM infrastructure. Requires proficiency in full-stack development, backend systems, cloud infrastructure, and production LLM experience.
155k – 340k/yr
HybridML Engineering
Software Engineer - E2E Autonomy
Applied IntuitionSunnyvale, CA
Builds ML tools, infrastructure, and manages large datasets for end-to-end autonomy research and productionizing self-driving software. Works with AI research and engineering teams to scale GPU compute, data, and evaluation systems. Requires strong software generalist skills across ML stack.
153k – 222k/yr
On-siteML Engineering
AI engineer
WriterSan Francisco, CA +1
Build and deploy scalable AI applications and intelligent agents for enterprise customers. Requires 5+ years of experience with Python, ML frameworks (PyTorch/TensorFlow/JAX), LLMs, and cloud platforms.
152k – 316k/yr
Hybrid5+ YOEML Engineering
Software Engineer, Forward Deployed Agent Builder
BrexSeattle, WA +2
Builds and deploys AI agents to automate internal workflows at Brex by embedding with teams, integrating systems and APIs, and measuring impact. Requires 4+ years experience shipping AI/automation with LLMs, agent frameworks, and databases.