Software Engineer, Developer Productivity, AI Tools
Thinking Machines LabSan Francisco, CA
Build and standardize AI-powered coding tools, agents, and dev environments to accelerate internal software development while maintaining security and quality. Requires experience with productivity tooling for large codebases, container/CI platforms, and AI model APIs.
350k – 475k/yrOn-site5+ YOEDevOps / SRE
Reliability Engineer, Supercomputing
Thinking Machines LabSan Francisco, CA
Ensure reliability of large GPU supercomputing clusters by diagnosing hardware/firmware/OS issues, automating monitoring, driving firmware rollouts, and working directly with vendors.
350k – 475k/yrOn-siteDevOps / SRE
Network Engineer, Supercomputing
Thinking Machines LabSan Francisco, CA
Own and debug multi-thousand-GPU network fabric (RDMA/RoCE, NVLink/NVSwitch) for large-scale AI training and inference. Requires backend language proficiency, large-scale cluster experience, and cross-stack ownership.
350k – 475k/yrOn-siteDevOps / SRE
Site Reliability Engineer (SRE)
Thinking Machines LabSan Francisco, CA
Site Reliability Engineer drives end-to-end reliability for AI fine-tuning platform Tinker, including SLOs, monitoring, incident response, and multi-tenant GPU scheduling. Requires distributed systems experience, software proficiency for reliability, and production incident handling.
350k – 475k/yrOn-siteDevOps / SRE
Research Engineer, Infrastructure, Training Systems
Thinking Machines LabSan Francisco, CA
Designs and optimizes distributed training systems scaling across thousands of GPUs for large AI models. Requires strong systems engineering, PyTorch/JAX expertise, and collaborative mindset to boost research productivity.
350k – 475k/yrOn-siteDevOps / SRE
Research Engineer, Infrastructure, RL Systems
Thinking Machines LabSan Francisco, CA
Designs and optimizes infrastructure for scalable reinforcement learning training of large models, improving reliability, observability, and throughput. Collaborates with researchers to productionize RL algorithms; requires strong engineering skills and deep learning framework knowledge.
350k – 475k/yrOn-siteDevOps / SRE
Research Engineer, Infrastructure, Numerics
Thinking Machines LabSan Francisco, CA
Designs and optimizes distributed training infrastructure for large-scale LLMs, focusing on low-precision numerics, kernel optimizations, and communication frameworks to enable stable, scalable trillion-parameter model training. Requires strong systems engineering, deep learning frameworks knowledge, and collaborative research mindset.
350k – 475k/yrOn-siteDevOps / SRE
Research Engineer, Infrastructure, Kernels
Thinking Machines LabSan Francisco, CA
Designs and optimizes high-performance ML kernels (CUDA, CuTe, Triton) for large-scale LLM training, focusing on GPU efficiency, low-precision formats, and distributed compute. Collaborates with researchers to bridge algorithms and hardware.
350k – 475k/yrOn-siteDevOps / SRE
Research Engineer, Infrastructure, Inference
Thinking Machines LabSan Francisco, CA
Designs, optimizes, and scales infrastructure for high-performance AI model inference, focusing on latency, throughput, efficiency, and reliability. Collaborates with researchers to enable production deployment of large-scale models using deep learning frameworks and distributed systems.
350k – 475k/yrOn-siteDevOps / SRE
Software Engineer, Supercomputing
Thinking Machines LabSan Francisco, CA
Designs, builds, and operates GPU supercomputing environments for large-scale AI training and inference. Automates cluster management, extends orchestration systems, and optimizes performance metrics in collaboration with researchers.
350k – 475k/yrOn-siteDevOps / SRE
Software Engineer, Systems Generalist
Thinking Machines LabSan Francisco, CA
Builds and scales core infrastructure for AI model training, data systems, and developer tools in a high-impact team. Requires backend proficiency (Python/Rust), experience with large-scale clusters like Kubernetes, and end-to-end project ownership.
350k – 475k/yrOn-siteDevOps / SRE
Search
Location
11 jobs
Job results
Software Engineer, Developer Productivity, AI Tools
Thinking Machines LabSan Francisco, CA
Build and standardize AI-powered coding tools, agents, and dev environments to accelerate internal software development while maintaining security and quality. Requires experience with productivity tooling for large codebases, container/CI platforms, and AI model APIs.
350k – 475k/yrOn-site5+ YOEDevOps / SRE
Reliability Engineer, Supercomputing
Thinking Machines LabSan Francisco, CA
Ensure reliability of large GPU supercomputing clusters by diagnosing hardware/firmware/OS issues, automating monitoring, driving firmware rollouts, and working directly with vendors.
350k – 475k/yrOn-siteDevOps / SRE
Network Engineer, Supercomputing
Thinking Machines LabSan Francisco, CA
Own and debug multi-thousand-GPU network fabric (RDMA/RoCE, NVLink/NVSwitch) for large-scale AI training and inference. Requires backend language proficiency, large-scale cluster experience, and cross-stack ownership.
350k – 475k/yrOn-siteDevOps / SRE
Site Reliability Engineer (SRE)
Thinking Machines LabSan Francisco, CA
Site Reliability Engineer drives end-to-end reliability for AI fine-tuning platform Tinker, including SLOs, monitoring, incident response, and multi-tenant GPU scheduling. Requires distributed systems experience, software proficiency for reliability, and production incident handling.
350k – 475k/yrOn-siteDevOps / SRE
Research Engineer, Infrastructure, Training Systems
Thinking Machines LabSan Francisco, CA
Designs and optimizes distributed training systems scaling across thousands of GPUs for large AI models. Requires strong systems engineering, PyTorch/JAX expertise, and collaborative mindset to boost research productivity.
350k – 475k/yrOn-siteDevOps / SRE
Research Engineer, Infrastructure, RL Systems
Thinking Machines LabSan Francisco, CA
Designs and optimizes infrastructure for scalable reinforcement learning training of large models, improving reliability, observability, and throughput. Collaborates with researchers to productionize RL algorithms; requires strong engineering skills and deep learning framework knowledge.
350k – 475k/yrOn-siteDevOps / SRE
Research Engineer, Infrastructure, Numerics
Thinking Machines LabSan Francisco, CA
Designs and optimizes distributed training infrastructure for large-scale LLMs, focusing on low-precision numerics, kernel optimizations, and communication frameworks to enable stable, scalable trillion-parameter model training. Requires strong systems engineering, deep learning frameworks knowledge, and collaborative research mindset.
350k – 475k/yrOn-siteDevOps / SRE
Research Engineer, Infrastructure, Kernels
Thinking Machines LabSan Francisco, CA
Designs and optimizes high-performance ML kernels (CUDA, CuTe, Triton) for large-scale LLM training, focusing on GPU efficiency, low-precision formats, and distributed compute. Collaborates with researchers to bridge algorithms and hardware.
350k – 475k/yrOn-siteDevOps / SRE
Get new-job notifications on iOS
Hotfix on iOS
Get a push summary when new jobs match your saved alerts.
Research Engineer, Infrastructure, Inference
Thinking Machines LabSan Francisco, CA
Designs, optimizes, and scales infrastructure for high-performance AI model inference, focusing on latency, throughput, efficiency, and reliability. Collaborates with researchers to enable production deployment of large-scale models using deep learning frameworks and distributed systems.
350k – 475k/yrOn-siteDevOps / SRE
Software Engineer, Supercomputing
Thinking Machines LabSan Francisco, CA
Designs, builds, and operates GPU supercomputing environments for large-scale AI training and inference. Automates cluster management, extends orchestration systems, and optimizes performance metrics in collaboration with researchers.
350k – 475k/yrOn-siteDevOps / SRE
Software Engineer, Systems Generalist
Thinking Machines LabSan Francisco, CA
Builds and scales core infrastructure for AI model training, data systems, and developer tools in a high-impact team. Requires backend proficiency (Python/Rust), experience with large-scale clusters like Kubernetes, and end-to-end project ownership.