SWE, ML
Build and own the software infrastructure supporting large-scale tabular model research, from experimentation through production. The role requires 5+ years of software engineering experience, expert Python and PyTorch skills, strong architecture expertise, and familiarity with modern ML tooling and cloud infrastructure.
About the job
Responsibilities
- Build and own the core codebase behind model development, making it fast to iterate on, robust for experiments and large training runs, and ready to carry models into production.
- Develop solutions, including agentic workflows, that improve researcher efficiency and effectiveness.
- Steward the long-term development and maintenance of the research codebase while preserving flexibility for empirical research.
- Lead pull request reviews and govern external contributions.
- Establish and implement software engineering best practices for machine learning.
- Mentor and guide research scientists on software engineering practices.
- Transition the core research codebase from an experimental state to a scalable, maintainable, and robust engineering standard.
- Manage technical debt, ensure experiment reproducibility, and improve overall code quality.
- Develop and maintain documentation for research infrastructure and systems.
- Collaborate across teams to integrate research code with other company systems and translate research breakthroughs into production impact.
- Build and own the testing layer alongside model development, including CI, contract tests, and interface verification.
Requirements
- 5+ years of software engineering experience, including meaningful experience with ML or data-intensive backends, training pipelines, distributed inference, feature infrastructure, or similar systems.
- Strong software architecture skills, including modular design, clean interfaces, and design patterns.
- Excellent communication and interdisciplinary collaboration skills.
- Expert-level Python proficiency and deep familiarity with PyTorch.
- Hands-on experience with modern ML research tooling and cloud infrastructure, such as AWS, GCP, Weights & Biases, Datadog, Kubernetes, ArgoCD, and GitHub Actions.
- Experience developing and using robustness and quality-assurance solutions for software or ML systems.
- Mission-driven mindset and a strong focus on software development excellence.
Nice to Have
- Familiarity with MLOps practices, CI/CD for machine learning, and infrastructure as code, such as Terraform.
- Experience in an early-stage, fast-paced startup environment.
- Background or strong interest in tabular data, data engineering, or enterprise data architecture.
- Experience with Rust, C++, or Go.
- Contributions to open-source machine learning libraries, frameworks, or tools.
- Bachelor's degree in Computer Science, Software Engineering, or a STEM field.
Compensation and Benefits
- Competitive compensation with salary and equity.
- Comprehensive health coverage for employees and dependents.
- Paid parental leave for all new parents, including adoptive and surrogate journeys.
- Relocation support for employees moving to an office location.
- Mission-driven, low-ego culture valuing diversity of thought, ownership, and bias toward action.
Skills
Python, PyTorch, AWS, GCP, Weights & Biases, Datadog, Kubernetes, Argo CD, GitHub Actions, MLOps, Terraform, Rust, C++, Go, CI/CD
Similar jobs
ML Engineering jobsBuild the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build and operate AI-powered, customer-facing workflows for Datadog Notebooks, combining reliable backend systems with LLM capabilities. The role requires 6+ years of engineering experience, Go or Python expertise, and experience delivering production AI products.
Build and deploy AI-powered features for conversation intelligence, developing production ML pipelines and inference services for voice and messaging data. The role requires 2+ years of applied ML experience, Python, an ML framework, NLP familiarity, and cloud infrastructure experience.
Build and operate edge MLOps infrastructure for smart-camera machine-learning systems, including model deployment, TensorRT compilation, fleet updates, telemetry, and reliability. The role requires production MLOps experience, embedded inference optimization, and strong collaboration with data-science and embedded-engineering teams.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.