Senior Member of Technical Staff, Synthetic Data
Build and optimize synthetic data and inference pipelines for large language models, combining research and software engineering to improve data quality, throughput, and model performance. The role requires strong Python and data-pipeline experience, familiarity with LLM inference frameworks, and experience with large-scale datasets.
About the job
Responsibilities
- Design and build scalable inference pipelines that run on large GPU clusters.
- Conduct data ablations to assess data quality and experiment with data mixtures to enhance model performance.
- Research and implement innovative synthetic data curation methods using Cohere’s infrastructure to advance natural language processing.
- Collaborate with researchers and engineers to ensure data pipelines meet the demands of cutting-edge language models.
Requirements
- Strong software engineering skills.
- Proficiency in Python and experience building data pipelines.
- Familiarity with data processing frameworks such as Apache Spark, Apache Beam, Pandas, or similar tools.
- Experience working with large language models through work projects, open-source contributions, or personal experimentation.
- Familiarity with LLM inference frameworks such as vLLM and TensorRT.
- Experience working with large-scale datasets, including web data, code data, and multilingual corpora.
- Passion for bridging research and engineering to solve complex data-related challenges in AI model training.
Nice to Have
- A paper at a top-tier venue such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP.
Compensation and Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Full health and dental benefits, including a separate mental health budget.
- RRSP matching, 401K, and pension scheme.
- 100% parental leave top-up for up to 6 months for either parent.
- Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation (30 working days).
- Travel budget for remote employees to visit other offices and an annual company offsite.
- Co-working benefit for employees not near an office.
- $500 home office stipend.
Skills
Python, Data Pipelines, Spark, Apache Beam, pandas, LLMs, vLLM, TensorRT, Gpu Clusters, Synthetic Data, Natural Language Processing, Web Data, Code Data, Multilingual Corpora, Data Ablation
Similar jobs
ML Engineering jobsBuild and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.
Build and deploy generative AI and LLM-powered agentic applications at Front to automate customer support inquiries, enhance product capabilities, and drive operational insights. Requires 5+ years software engineering experience with strong production AI/ML focus, agentic/RAG expertise, and proficiency in Node.js, TS, and Python.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.
Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.