Skip to content
Thinking Machines LabThinking Machines LabSan Francisco, CA

Software Engineer, Developer Productivity, AI Tools

Build and standardize AI-powered coding tools, agents, and dev environments to accelerate internal software development while maintaining security and quality. Requires experience with productivity tooling for large codebases, container/CI platforms, and AI model APIs.

350k – 475k/yr
On-site5+ YOEDevOps / SRE

About the role

What You’ll Do

  • Enable our researchers and engineers to leverage AI to improve coding productivity without compromising code quality
  • Standardize AI coding tools, such as Claude Code, Cursor, and Codex. Help configure, harden, and maintain the best tools, integrating org-wide configurations with individual preferences.
  • Build secure, reproducible agent sandboxes for remote dev & CI testing.
  • Set up golden-path dev environments and guardrails for secrets/PII.
  • Help individual contributors develop their personalized AI-enabled workflow.
  • Track tool usage, reliability, and cost.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent industry experience in computer science, engineering, or similar.
  • Experience developing productivity tools and best practices for large codebases.
  • Ability to communicate clearly and work with researchers to build and manage a variety of internal tools.

Preferred qualifications:

  • Hands-on experience with container platforms (e.g. Docker/Kubernetes), modern CI (GitHub Actions/Buildkite), and package management tools (uv).
  • Practical experience with AI coding tools and model APIs (e.g. OSS via vLLM / SGLang / TGI).
  • Solid Linux/networking fundamentals; comfort with secrets management and safe egress.
  • Proficiency in systems programming languages (e.g. Rust) and scripting languages (e.g. Python).

Skills

DockerKubernetesGitHub ActionsbuildkiteuvvLLMsglangtgiLinuxRustPythonai coding toolssecrets management

Similar roles

DevOps / SRE jobs
Thinking Machines Lab

Reliability Engineer, Supercomputing

Thinking Machines LabSan Francisco, CA

Ensure reliability of large GPU supercomputing clusters by diagnosing hardware/firmware/OS issues, automating monitoring, driving firmware rollouts, and working directly with vendors.

350k – 475k/yr
On-siteDevOps / SRE
Thinking Machines Lab

Network Engineer, Supercomputing

Thinking Machines LabSan Francisco, CA

Own and debug multi-thousand-GPU network fabric (RDMA/RoCE, NVLink/NVSwitch) for large-scale AI training and inference. Requires backend language proficiency, large-scale cluster experience, and cross-stack ownership.

350k – 475k/yr
On-siteDevOps / SRE
Anthropic

Performance Engineer, Inference Systems

AnthropicSan Francisco, CA +2

Performance engineer focused on cross-layer investigations of Anthropic's inference fleet for Claude, optimizing throughput, latency, reliability, and correctness while building observability and partnering with kernel and serving teams.

350k – 850k/yr
HybridDevOps / SRE
Thinking Machines Lab

Site Reliability Engineer (SRE)

Thinking Machines LabSan Francisco, CA

Site Reliability Engineer drives end-to-end reliability for AI fine-tuning platform Tinker, including SLOs, monitoring, incident response, and multi-tenant GPU scheduling. Requires distributed systems experience, software proficiency for reliability, and production incident handling.

350k – 475k/yr
On-siteDevOps / SRE
Thinking Machines Lab

Research Engineer, Infrastructure, Training Systems

Thinking Machines LabSan Francisco, CA

Designs and optimizes distributed training systems scaling across thousands of GPUs for large AI models. Requires strong systems engineering, PyTorch/JAX expertise, and collaborative mindset to boost research productivity.

350k – 475k/yr
On-siteDevOps / SRE