Skip to content
MongoDBMongoDBUnited States

Software Engineer 3

Build and maintain tooling, evaluation systems, quality gates, and infrastructure for MongoDB's agent skills and AI platform. Requires 2+ years building production software, developer tools, CLIs, test infrastructure, or CI/CD, with strong fundamentals in API design, testing, and reasoning about nondeterministic AI systems.

109k – 215k/yr
Remote2+ YOEML Engineering

About the role

What you'll do

  • Build and maintain agent skills and the infrastructure to validate, evaluate, publish, and maintain them
  • Design evaluation datasets and workflows that compare agent behavior against a baseline and produce actionable quality signals
  • Build agent metrics and observability: skill selection and routing, success and failure outcomes, tool calls, latency, and token usage
  • Design safety and quality gates for agent-authored content: rule packs, static analysis, confidence thresholds, structured verdicts, and bounded suppression
  • Create CLIs, libraries, and MCP integrations that other repositories adopt and that run in local development and CI
  • Integrate tooling into GitHub Actions and other CI workflows, including secrets, annotations, exit codes, and artifacts
  • Build code-generation quality checks, such as anti-pattern catalogs and linting for AI-generated MongoDB code
  • Investigate real failures such as nondeterministic results, false positives, and unsafe generated guidance, and turn them into reusable improvements
  • Collaborate with engineers, security partners, and product teams; communicate trade-offs, risks, and ownership across teams

Examples of the problems you'll solve

  • How can tests verify an agent tool's structured result when item order may vary, but counts, required fields, and values must remain correct
  • How can a CI gate flag unsafe instructions in an agent skill without treating every neutral mention as an incident or letting cautionary wording hide a real instruction
  • How can an evaluation suite show whether a skill improves answers over a baseline and give authors enough signal to improve it

What we're looking for

  • 2+ years of experience building production software, developer tools, internal platforms, or automation systems
  • Software engineering fundamentals in API design, testing, error handling, and maintainability
  • Experience building CLIs, libraries, test infrastructure, static analysis, or CI/CD workflows
  • Ability to design systems that are usable by developers and reliable in automation
  • Experience reasoning about correctness and safety with ambiguous input, nondeterministic output, false positives, or untrusted content
  • Comfort in an evolving R&D environment where the right abstraction emerges through prototypes and feedback
  • Written and verbal communication, including explaining technical trade-offs and aligning stakeholders across teams

Nice to have

  • Experience with agentic systems, LLM applications, prompt or rubric-based evaluation, or AI-assisted development
  • Experience building eval harnesses, benchmark datasets, quality metrics, LLM-as-judge workflows, or human-review tooling
  • Experience with Go, Python, JavaScript/TypeScript, Java, or C#
  • Experience with GitHub Actions security, secret handling, static rule engines, or policy enforcement
  • Experience moving prototypes into production

What success looks like

In your first year, you will:

  • Ship tooling that makes agent skills or developer workflows easier to test, review, and adopt
  • Improve the quality and interpretability of evaluations, not just their count
  • Convert recurring manual work and fragile scripts into documented, reusable automation
  • Make security, correctness, and operational trade-offs explicit in the designs you ship
  • Earn adoption from partner teams through clear interfaces and reliable CI
  • Own projects independently while collaborating on shared systems

Skills

cliCI/CDGitHub Actionsstatic analysisPythonGoJavaScriptTypeScriptJavaC#LLMsAPI Designtesting

Similar roles

ML Engineering jobs
Topaz Labs

Software Engineer, AI Inference / HPC

Topaz LabsDallas, TX

Develops and optimizes AI inference engine for image/video enhancement, focusing on performance, GPU/CPU optimization, model deployment, and hardware partnerships. Requires C/C++ expertise, 1+ years experience in performance optimization and image processing.

110k – 150k/yrOn-site1+ YOEML Engineering
Pinterest

Data Scientist II, ML Infrastructure

PinterestPalo Alto, CA

Build and productionize ML measurement, causal inference, and platform tooling at Pinterest. Translate research into scalable pipelines, develop self-serve causal tools, and create centralized systems for feature importance, model evaluation, and infrastructure efficiency.

114k – 235k/yrHybrid2+ YOEML Engineering
Deepgram

Software Engineering

DeepgramCalifornia

Intern on the engineering team owning and shipping a real scoped project end-to-end (design to production code) in voice AI, ASR, TTS or related systems. Requires self-motivated builder with first-principles reasoning, AI-first mindset, and ability to quickly learn new languages/codebases.

114k – 135k/yrRemoteEntry levelML Engineering
Wonderschool

Early Career Software Engineer – Applied AI

WonderschoolSan Francisco, CA

Early career software engineer builds and integrates AI agents and solutions using frameworks like LangChain and RAG pipelines to enhance childcare platform features. Requires bachelor's in CS/engineering, Python/JS proficiency, and AI/ML familiarity; hybrid onsite in SF office 3 days/week.

100k – 120k/yrHybridEntry levelML Engineering
Ai2

Research Engineer, Asta

Ai2Seattle, WA

Research Engineer builds ML infrastructure and agentic systems to accelerate scientific discovery in biology, neuroscience, and more. Requires 2+ years experience with Python, PyTorch/JAX/TF, cloud resources, and deep learning for LLM/agent research at Ai2.

119k – 178k/yrOn-site2+ YOEML Engineering