Senior Machine Learning Engineer - Model Evaluations, Public Sector
Senior Machine Learning Engineer building and deploying production GenAI, agentic systems, LLMs, and computer vision models for mission-critical public sector and defense applications at Scale AI. Requires active security clearance, extensive production ML experience, and strong Python/TF/PyTorch skills.
About the job
Responsibilities
- Take state-of-the-art models developed internally and from the community, and use them in production to solve problems for customers and taskers.
- Improve and maintain production models through retraining, hyperparameter tuning, and architectural updates, while preserving core performance characteristics.
- Collaborate with product and research teams to identify and prototype ML-driven product enhancements, including for upcoming product lines.
- Work with massive datasets to develop both generic models as well as fine-tune models for specific products.
- Build scalable machine learning infrastructure to automate and optimize our ML services.
- Serve as a cross-functional representative and advocate for machine learning techniques across engineering and product organizations.
- Learn new technologies quickly and manage multiple priorities in a fast-paced environment.
- Travel lightly (approximately 10%) for customer interaction and team needs.
Requirements
- Active security clearance.
- Extensive experience with GenAI, Agentic AI, natural language processing, deep learning and deep reinforcement learning, or computer vision in a production environment.
- Solid background in algorithms, data structures, and object-oriented programming.
- Strong programming skills in Python, experience in TensorFlow or PyTorch.
Nice-to-Haves
- Graduate degree in Computer Science, Machine Learning or Artificial Intelligence specialization.
- Experience working with cloud platforms (e.g. AWS or GCP) and deploying machine learning models in cloud environments.
- Experience with computer vision, generative AI models, large language models, or agentic systems.
- Familiarity with ML evaluation frameworks and agentic model design.
Compensation and Benefits
- Base salary range for this full-time position in Washington, DC: $225,750–$282,450.
- Compensation packages include base salary, equity, and benefits.
- Benefits include comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. May be eligible for additional benefits such as a commuter stipend.
- Equity grant subject to Board of Director approval.
Skills
Generative AI, Agentic AI, Natural Language Processing, Deep Learning, Deep Reinforcement Learning, Computer Vision, Python, TensorFlow, PyTorch, AWS, GCP
Similar jobs
ML Engineering jobsLeads applied research for next-generation fraud detection models across graph, sequential, image, and video data, translating prototypes into production solutions. Requires strong Python skills, research leadership, and a PhD or equivalent research experience.
Senior Machine Learning Engineer building and scaling production fraud-detection systems, including feature pipelines, online inference, monitoring, and reliability capabilities. Requires 6+ years of experience with ML infrastructure and technologies such as Python, PyTorch, Spark, SageMaker, and Airflow.
Develop and deploy machine learning models that detect fraud, from feature engineering and experimentation through production optimization. The role requires 7+ years of ML or software engineering experience, strong Python and SQL skills, and expertise in evaluating and scaling reliable models.
Build and deploy machine learning products for a new consumer-facing financial application, from opportunity discovery and experimentation through production and ongoing optimization. The role requires 6+ years of ML experience, strong Python and SQL skills, and the ability to work across product and engineering teams.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.