# Software Engineer 3

**Company:** [MongoDB](https://hotfix.jobs/companies/mongodb)
**Location:** Remote
**Role:** ML Engineering
**Salary:** $109k – $215k/yr
**Experience:** 2+ years
**Skills:** cli, CI/CD, GitHub Actions, static analysis, Python, Go, JavaScript, TypeScript, Java, C#, LLMs, API Design, testing
**Posted:** 2026-08-12

> Build and maintain tooling, evaluation systems, quality gates, and infrastructure for MongoDB's agent skills and AI platform. Requires 2+ years building production software, developer tools, CLIs, test infrastructure, or CI/CD, with strong fundamentals in API design, testing, and reasoning about nondeterministic AI systems.

## Job Description

## What you'll do

- Build and maintain agent skills and the infrastructure to validate, evaluate, publish, and maintain them
- Design evaluation datasets and workflows that compare agent behavior against a baseline and produce actionable quality signals
- Build agent metrics and observability: skill selection and routing, success and failure outcomes, tool calls, latency, and token usage
- Design safety and quality gates for agent-authored content: rule packs, static analysis, confidence thresholds, structured verdicts, and bounded suppression
- Create CLIs, libraries, and MCP integrations that other repositories adopt and that run in local development and CI
- Integrate tooling into GitHub Actions and other CI workflows, including secrets, annotations, exit codes, and artifacts
- Build code-generation quality checks, such as anti-pattern catalogs and linting for AI-generated MongoDB code
- Investigate real failures such as nondeterministic results, false positives, and unsafe generated guidance, and turn them into reusable improvements
- Collaborate with engineers, security partners, and product teams; communicate trade-offs, risks, and ownership across teams

## Examples of the problems you'll solve

- How can tests verify an agent tool's structured result when item order may vary, but counts, required fields, and values must remain correct
- How can a CI gate flag unsafe instructions in an agent skill without treating every neutral mention as an incident or letting cautionary wording hide a real instruction
- How can an evaluation suite show whether a skill improves answers over a baseline and give authors enough signal to improve it

## What we're looking for

- 2+ years of experience building production software, developer tools, internal platforms, or automation systems
- Software engineering fundamentals in API design, testing, error handling, and maintainability
- Experience building CLIs, libraries, test infrastructure, static analysis, or CI/CD workflows
- Ability to design systems that are usable by developers and reliable in automation
- Experience reasoning about correctness and safety with ambiguous input, nondeterministic output, false positives, or untrusted content
- Comfort in an evolving R&D environment where the right abstraction emerges through prototypes and feedback
- Written and verbal communication, including explaining technical trade-offs and aligning stakeholders across teams

## Nice to have

- Experience with agentic systems, LLM applications, prompt or rubric-based evaluation, or AI-assisted development
- Experience building eval harnesses, benchmark datasets, quality metrics, LLM-as-judge workflows, or human-review tooling
- Experience with Go, Python, JavaScript/TypeScript, Java, or C#
- Experience with GitHub Actions security, secret handling, static rule engines, or policy enforcement
- Experience moving prototypes into production

## What success looks like

In your first year, you will:

- Ship tooling that makes agent skills or developer workflows easier to test, review, and adopt
- Improve the quality and interpretability of evaluations, not just their count
- Convert recurring manual work and fragile scripts into documented, reusable automation
- Make security, correctness, and operational trade-offs explicit in the designs you ship
- Earn adoption from partner teams through clear interfaces and reliable CI
- Own projects independently while collaborating on shared systems

## Similar roles

- [Software Engineer, AI Inference / HPC](https://hotfix.jobs/jobs/77dc9fec-7fe8-426d-9a36-e19ef1ea6fd3) - Topaz Labs - Dallas, TX - $110k – $150k/yr
- [Data Scientist II, ML Infrastructure](https://hotfix.jobs/jobs/ca4fe289-b334-4cec-85e3-3024d1e3762c) - Pinterest - Palo Alto, CA - $114k – $235k/yr
- [Software Engineering](https://hotfix.jobs/jobs/fd8fddd7-ecf7-427f-b30c-b406ca715739) - Deepgram - Remote - $114k – $135k/yr
- [Early Career Software Engineer – Applied AI](https://hotfix.jobs/jobs/a86c7eb4-2803-4e64-9563-042b860574d5) - Wonderschool - San Francisco, CA - $100k – $120k/yr
- [Research Engineer, Asta](https://hotfix.jobs/jobs/50766069-cf57-4073-b1bd-a9b997b92021) - Ai2 - Seattle, WA - $119k – $178k/yr

**Apply:** https://hotfix.jobs/jobs/ffa1abd6-cf7d-4412-95c2-7350792135b4
**Canonical:** https://hotfix.jobs/jobs/ffa1abd6-cf7d-4412-95c2-7350792135b4