Senior Data Science Engineer
Owns the full lifecycle of data and ML solutions, from ingestion and feature-ready datasets through production deployment and business-impact measurement. The role combines data engineering, applied machine learning, MLOps, and generative AI to build risk detection capabilities.
About the job
Responsibilities
Data Science & Applied ML
- Research, prototype, and develop machine learning and LLM-based models for complex business problems, including risk detection and prioritization.
- Wrap models in production-ready APIs and integrate them into the core product.
- Make model outputs interpretable by translating predictions into actionable reason codes.
- Partner with operational teams to gather feedback, refine features, and improve model relevance.
Data Engineering
- Design, build, and maintain scalable pipelines that ingest disparate data sources into a data warehouse or lake.
- Implement data validation, quality checks, and transformation workflows across raw, curated, and serving layers.
- Build and maintain curated datasets for analytics and model training.
MLOps & Production Ownership
- Implement and maintain CI/CD pipelines for data workflows and ML model deployment across environments.
- Monitor pipeline latency, data drift, and model performance in production; design alerting and retraining triggers.
- Define success metrics, track ROI, and iterate based on real-world model efficacy.
- Manage infrastructure as code and containerized deployments for reproducible releases.
Requirements
- 5–8+ years spanning data engineering and data science/ML, with a track record of shipping models to production.
- Strong Python proficiency.
- Experience with Spark or PySpark for large-scale data processing.
- Advanced SQL for complex transformation, analysis, and data modeling.
- Experience with cloud data platforms such as Databricks or Snowflake.
- Experience with ETL/ELT frameworks such as dbt, Lakeflow Declarative Pipelines, Databricks Autoloader, Informatica, or similar.
- Familiarity with ML experiment tracking tools such as MLflow or Weights & Biases.
- Git-based development, branching strategies, CI/CD, infrastructure as code, and Docker.
- Experience with orchestration tools such as Databricks Workflows or Apache Airflow.
Nice-to-Haves
- Production experience with LLMs and generative AI techniques, including prompt engineering, RAG architectures, fine-tuning, or evaluation frameworks.
- Experience building or operating ML platforms, feature stores, or model registries.
- Experience in risk, compliance, fraud detection, or other high-stakes ML domains.
Compensation & Benefits
- Competitive compensation.
- Flexible work options.
- Visa sponsorship is not available.
- International remote work is not supported.
Skills
Python, Spark, Pyspark, SQL, Databricks, Snowflake, dbt, MLflow, Weights & Biases, Git, CI/CD, Terraform, Docker, Apache Airflow, LLMs
Similar jobs
ML Engineering jobsLeads hands-on development and deployment of production AI agents and workflows, sets technical standards, and mentors engineers. Requires 4+ years of software engineering experience, production LLM application experience, and strong Python or TypeScript skills.
Build and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
Develop and deploy real-time perception and sensor-fusion software for autonomous battery-electric rail vehicles. The role requires strong robotics, geometry-based computer vision, C/C++ and Rust experience, plus hands-on work with multimodal sensors and production systems.
Build and optimize large language model training and post-training pipelines, improving model quality, distributed performance, evaluation, and production readiness. The role requires deep PyTorch and transformer experience, strong distributed-systems and software-engineering skills, and expertise in modern LLM optimization techniques.
Build and operate backend infrastructure for machine learning model training, serving, feature management, and marketplace simulation. The role requires 6+ years of software engineering experience, distributed systems expertise, and experience with production ML platforms.