Skip to content
IntercomIntercom

AI Infrastructure Engineer

Build and optimize large-scale training pipelines and low-latency inference services for Fin’s AI models. The role requires strong software engineering experience, hands-on expertise in model training, inference, or GPU programming, and collaboration with ML scientists.

About the job

Responsibilities

  • Implement and scale training pipelines for large transformer and LLM models, from data ingestion and preprocessing through distributed training and evaluation.
  • Build and optimize inference services that deliver low-latency, high-reliability customer experiences, including autoscaling, routing, and fallbacks.
  • Tune GPU kernels, improve utilization, and identify bottlenecks across the training and inference stack.
  • Collaborate with ML scientists to implement advanced training and inference methods and bring them to production.
  • Participate in hiring, mentoring, and developing engineers.
  • Raise technical standards, reliability, and operational excellence across the AI platform.

Requirements

  • 5+ years of software engineering experience with a strong record of shipping high-quality products or platforms.
  • Degree in Computer Science, Computer Engineering, or a related field, or equivalent experience with strong fundamentals.
  • Hands-on experience with model training, particularly transformers and LLMs; model inference at scale; or low-level GPU work such as CUDA or Triton kernels.
  • Experience working in production environments at meaningful traffic, data, or organizational scale.
  • Clear communication and ability to explain complex technical topics to technical and non-technical audiences.
  • Strong technical fundamentals and willingness to learn and develop.
  • Deep knowledge of at least one programming language, such as Python, Ruby, Java, or Go.

Nice-to-haves

  • Experience at AI-native companies training or serving their own models.
  • Experience running training or inference workloads on Kubernetes.
  • Experience with AWS or other major cloud providers.
  • Production Python experience in ML or infrastructure contexts.
  • Personal projects, open-source contributions, meetups, or published technical content.

Compensation & Benefits

  • Competitive salary and equity.
  • Lunch on weekdays, snacks, and a stocked kitchen.
  • Regular compensation reviews.
  • Unlimited access to Claude Code and other AI tools.
  • Pension scheme with matching up to 4%.
  • Life assurance and comprehensive health and dental insurance for employees and dependents.
  • Flexible paid time off.
  • Paid maternity leave and six weeks of paternity leave.
  • Cycle-to-Work Scheme and secure bike storage.
  • MacBooks as standard, with Windows available for certain roles.
  • Hybrid working policy requiring at least three days per week in the office.

Skills

Python, Ruby, Java, Go, CUDA, Triton, Kubernetes, AWS, Transformers, LLMs, Distributed Training, Model Inference, Gpu Optimization, Autoscaling, Data Preprocessing

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

OpenAI

OpenAI

London, United Kingdom

Applied AI Engineer, Digital Natives
No salary listedHybridML Engineering

Build and deploy AI-powered products for digital-native customers, taking systems from experimentation through production and scale. The role requires strong Python skills, hands-on production engineering, systematic AI evaluation, and the ability to navigate reliability, security, governance, and customer impact.

Elliptic

Elliptic

London, United Kingdom

Agent Engineer
No salary listedHybrid5+ YOEML Engineering

Build full-stack AI agent fleets, APIs, workflows, and internal services that automate complex business processes. The role requires at least five years of engineering experience, hands-on LLM framework experience, production AWS expertise, Kubernetes, and strong API and database skills.

Protege

Protege

Remote

AI Engineer - New Verticals
No salary listedRemote3+ YOEML Engineering

Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.

Build

Build

New York, NY
AI Engineer - Assistant Experience
$120k+/yrOn-siteML Engineering

Build and operate Dougie, an agentic AI system that executes workflows, evaluates its own performance, retains institutional context, and improves in production. The role requires experience deploying unattended agentic systems and engineering reliable memory, retrieval, orchestration, and feedback loops.