Data Scientist - Extensions
Research and develop data science methods to enhance NEXUS tabular model performance across enterprise datasets. Design production-grade Python components and run rigorous experiments to improve prediction on real-world structured data.
About the job
Key Responsibilities
- Research and develop data science methods that improve NEXUS predictive performance across diverse enterprise datasets, industries, and prediction task types
- Design and implement robust, production-quality Python components with a strong focus on correctness, generality, and reusability
- Deeply understand the characteristics of real-world enterprise data and develop strategies that help NEXUS handle them reliably
- Run rigorous experiments to measure the impact of new approaches, design meaningful benchmarks, and use results to guide prioritization
- Work across a wide variety of structured data problems - including but not limited to classification, regression, ranking, and forecasting
- Collaborate closely with the Engineering and Research teams to develop a deep understanding of NEXUS model behavior and use that knowledge to inform your work
- Work with Applied AI Engineers to validate approaches on real customer datasets and translate findings into product capabilities
- Contribute to technical documentation and internal best practices, helping the broader team apply new capabilities correctly and confidently
Must Have
- 5+ years of experience in data science or machine learning roles
- Strong Python skills, including fluency with pandas, numpy, and scikit-learn
- Deep hands-on experience with traditional ML models: XGBoost, LightGBM, CatBoost, and similar gradient boosting frameworks
- Solid understanding of what makes real-world tabular data challenging: class imbalance, high cardinality, distribution shift, missing values, and more
- Strong experimental mindset - comfortable designing benchmarks and drawing rigorous conclusions from noisy results
- Ability to work autonomously and drive work from idea to shipped output
Nice to Have
- Familiarity with tabular foundation models (TabPFN, CARTE, or similar)
- Competitive data science experience (Kaggle, DrivenData, or similar) - especially top finishes on tabular competitions
- Background in a domain where structured prediction matters: finance, supply chain, healthcare, retail, or industrial
- Experience contributing to or designing internal ML libraries or shared tooling
- Familiarity with DuckDB, Polars, or modern in-process analytics engines
- Comfort reading ML research papers and translating findings into practical implementations
Benefits
- Competitive compensation with salary and equity
- Comprehensive health coverage for you and your dependents
- Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys
- Relocation support for employees moving to join the team in one of our office locations
- A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action
Skills
Python, pandas, NumPy, scikit-learn, Xgboost, Lightgbm, Catboost, Tabpfn, Duckdb, Polars
Similar jobs
Data Science jobsThe Senior Data Scientist will build scalable product analytics assets, measurement frameworks, and AI-assisted analytical workflows while partnering with cross-functional teams to guide product and business decisions. The role requires 4+ years of applied data experience and strong SQL, programming, statistics, machine learning, and experimentation skills.
Staff Data Scientist driving search ranking and results improvements through machine learning, NLP, evaluation, and ranking-signal development. Requires 10+ years of data science experience, staff-level leadership, production ML deployment, Python, and advanced SQL skills.