Skip to content
Beacon AIBeacon AISan Carlos, CA

Software Engineer, Artificial Intelligence/LLM

Build and ship production LLM-powered features for aviation safety and efficiency, including RAG, tool-calling, evals, guardrails, and monitoring for cost/latency/quality. Requires prior shipped LLM applications, strong production coding, and RAG depth; hybrid in San Carlos with multiple seniority levels available.

135k – 260k/yr
Hybrid5+ YOEML Engineering

About the role

What you'll do

  • Build user-facing LLM features: Design and implement retrieval-augmented generation and tool-calling flows using frameworks like LangChain (prefer simpler primitives). Deliver robust JSON/schema-bound outputs with validation, retries, and fallbacks. Add function calling to integrate with internal tools, search, routing, and data services.
  • Own the service layer: Ship APIs and workers in Python or TypeScript with clear contracts, streaming, and backoff. Add caching, request shaping, prompt templates, and context packing to control latency and cost. Integrate with AWS Bedrock, OpenAI, Anthropic, or self-hosted endpoints.
  • Retrieval and data prep: Collaborate on chunking, embeddings, and indexing for documents, time series, and multimedia. Choose and tune vector backends (OpenSearch, pgvector, Pinecone). Keep knowledge bases fresh via syncs from S3, Aurora, DynamoDB, and external sources.
  • Evaluation and quality: Create offline evals and golden sets for prompts, retrievers, and tools. Stand up online metrics for task success, hallucination rate, retrieval precision/recall, p95 latency, and cost per request. Run A/B tests and prompt/version rollouts with guardrails and canaries.
  • Safety, privacy, and compliance: Implement content/policy checks, PII detection/redaction, access controls, and auditing. Design human-in-the-loop paths. Handle aviation data per internal security standards.
  • Operate what you build: Add tracing, logs, and dashboards for model calls, token usage, errors, and saturation. Debug failures across retrieval, prompts, tools, and providers.

Requirements

  • Shipped LLM apps in production and improved them with data.
  • Strong builder: comfortable writing production code, tests, and docs; keep things simple and observable.
  • Deep RAG and tools experience: understand embeddings, chunking, vector search tradeoffs, and function calling.
  • Quality mindset: design evals, define success metrics, iterate on evidence.
  • Cost and latency aware: track p95, hit SLAs, reduce cost without sacrificing quality.
  • Clear communicator: explain tradeoffs and align with product, infra, and security partners.

Nice-to-haves

  • Experience with Bedrock, OpenSearch Serverless, pgvector, Pinecone, or Weaviate.
  • Prompt versioning, guardrails, and provider routing in production.
  • Multimodal work with time series or video.
  • Familiarity with GPU inference, Triton, or TensorRT-LLM.
  • Aviation or safety-critical domain exposure.
  • DevOps basics (CI/CD, IaC, secure secrets handling).

Compensation & Benefits

  • Salary: $135,000 - $260,000 (varies by level)
  • Healthcare: 100% employee medical premiums covered; 25% for dependents
  • Time off: 3 weeks PTO + 13+ paid company holidays
  • Monthly phone and wellness stipends
  • 401(k) offered

Skills

LLMsRAGLangChainPythonTypeScriptaws bedrockOpenAIAnthropicEmbeddingsvector searchopensearchpgvectorpineconeEvaluation FrameworksPrompt Engineering

Similar roles

ML Engineering jobs
Beacon AI

Software Engineer, Cloud Infrastructure

Beacon AISan Carlos, CA

Build and operate scalable AWS cloud infrastructure powering LLM platforms, RAG systems, data pipelines, and IoT for aviation AI applications. Requires strong AWS depth, LLM/RAG production experience, Python data engineering skills, and end-to-end ownership.

135k – 260k/yrHybrid5+ YOEML Engineering
Assembled

Software Engineer - Forecasting & Scheduling

AssembledUnited States

Builds forecasting interfaces, data pipelines, and scheduling systems for thousands of support agents using ML models. Requires Python ML libraries experience and background in ML/algorithmic teams with focus on performance optimization.

135k – 280k/yrRemoteML Engineering
Assembled

Software Engineer - Forecasting & Scheduling

AssembledSan Francisco, CA

Develops forecasting interfaces, data pipelines, and scheduling systems to predict support contact volume and optimize agent schedules for thousands, incorporating ML models and constraints like labor laws. Requires Python/ML experience and performance focus.

135k – 280k/yrHybridML Engineering
Assembled

Software Engineer - AI Agents & Platform

AssembledSan Francisco, CA

Builds autonomous AI agents and platform infrastructure for customer support, enhancing LLM performance with RAG techniques, scaling systems with Golang, and integrating STT/TTS. Requires 5+ years software engineering experience with LLMs.

135k – 280k/yrOn-site5+ YOEML Engineering
Virta Health

AI Operations Engineer

Virta HealthUnited States

Serve as technical SME driving adoption of agentic AI platform across departments. Build AI observability dashboards, govern vendor risks, integrate AI systems, and create repeatable IT workflows for safe, scalable generative AI use.

136k – 176k/yrRemote4+ YOEML Engineering