Staff Engineer responsible for designing enterprise AI platform architecture, reusable components, responsible AI controls, observability, and governance patterns. Requires deep expertise in generative AI, RAG, agentic workflows, and influencing cross-functional teams without direct authority.
190k – 290k/yr
Remote7+ YOEML Engineering
About the role
What you'll do
AI Platform Architecture & Standards
Define and evolve enterprise AI architecture patterns for LLM integration, retrieval-augmented generation, agentic workflows, prompt orchestration, and workflow automation.
Create reference architectures, design reviews, decision records, and implementation guidance that enable consistent AI development across business units.
Serve as a technical authority for AI platform decisions, including model selection, integration approaches, data boundary enforcement, and lifecycle management.
Evaluate emerging AI technologies and recommend fit-for-purpose adoption paths aligned to security, operational, and enterprise architecture requirements.
Partner with product, platform, and business technology teams to identify common needs and convert them into reusable engineering patterns.
Reusable Components & Shared Services
Design and build reusable AI components such as connectors, agents, skill templates, prompt libraries, data pipelines, integration adapters, and service APIs.
Lead technical design for shared platform services for AI observability, logging, usage metering, evaluation, and lifecycle management.
Establish quality, versioning, deprecation, documentation, and contribution standards for the shared AI component catalog.
Guide teams through adoption of shared components, balancing standardization with practical implementation needs.
Identify opportunities to eliminate duplicate AI engineering efforts through consolidation, abstractions, and platformization.
Responsible AI Engineering & Governance
Architect engineering controls for access management, data classification enforcement, prompt safety, output validation, audit logging, and policy adherence.
Partner with Security, Legal, and compliance stakeholders to embed responsible AI requirements into development and deployment pipelines.
Design model and agent lifecycle governance patterns, including version tracking, evaluation, drift monitoring, rollback, and deprecation workflows.
Build technical dashboards and telemetry that expose adoption, risk, performance, and governance compliance across AI-enabled systems.
Represent engineering considerations in AI governance reviews and translate policy requirements into implementable technical standards.
Productivity, Measurement & Technical Leadership
Develop AI-assisted workflow patterns that improve individual productivity, team collaboration, knowledge retrieval, meeting intelligence, document generation, and task automation.
Design measurement approaches that connect AI usage to time savings, quality improvement, error reduction, capacity creation, and business value.
Partner with Finance and platform teams to develop cost metering, showback/chargeback, and optimization mechanisms for AI services.
Mentor senior and mid-level engineers, raise engineering quality, and lead complex cross-functional technical initiatives from concept through production.
Contribute to communities of practice, internal enablement material, and technical evangelism for enterprise AI engineering standards.
Required qualifications
Progressive experience in enterprise software engineering, AI platform engineering, data platform engineering, or digital workplace technology roles.
Deep hands-on understanding of generative AI, large language model integration, RAG architectures, agentic AI patterns, prompt orchestration, and production AI system design.
Experience designing shared platform services, reusable component libraries, APIs, integration frameworks, or developer enablement platforms used by multiple teams.
Strong architecture judgment across security, reliability, scalability, observability, maintainability, and operational cost tradeoffs.
Experience implementing or contributing to AI governance controls such as access management, data classification, audit logging, model lifecycle management, and compliance-aware development practices.
Ability to influence technical direction across matrixed teams through architecture reviews, written guidance, reference implementations, and hands-on collaboration.
Experience defining metrics, telemetry, or attribution mechanisms for adoption, productivity, cost, quality, or operational performance.
Strong written and verbal communication skills with the ability to explain complex AI engineering concepts to technical and non-technical audiences.
Preferred qualifications
Experience in regulated, defense-adjacent, security-sensitive, or data-governed environments.
Background in MLOps, AI observability, model evaluation frameworks, agent evaluation, and production monitoring.
Familiarity with enterprise AI tooling ecosystems including copilot platforms, workflow automation suites, vector databases, and enterprise search/RAG platforms.
Hands-on experience with enterprise data platforms such as Databricks, Snowflake, lakehouse architectures, or comparable data foundations.
Experience implementing usage metering, cost allocation, showback/chargeback, or AI spend optimization capabilities.
Track record of mentoring engineers and raising technical standards without relying on direct management authority.
Advanced degree in Computer Science, Engineering, Data Science, or a related technical field.
Build and optimize LLM inference infrastructure at enterprise scale for partner and self-hosted frontier models. Requires 8+ years backend/infrastructure engineering experience with distributed systems, real-time serving, and ML/GPU orchestration.
190k – 265k/yr
On-site8+ YOEML Engineering
Staff Software Engineer - AI Research Infrastructure
DatabricksNew York, NY
Founding member of a new team building foundational evaluation infrastructure and flywheels for Databricks' AI/Genie Agents. Design scalable tooling for benchmarking, regression detection, and quality measurement that drives continuous agent improvement across research, training, and production.
190k – 270k/yr
On-site6+ YOEML Engineering
Staff Software Engineer, Foundation Model API
DatabricksSan Francisco, CA
Build and shape the Foundation Model API serving layer for large-scale LLM inference (partner and self-hosted models) at Databricks. Requires 8+ years backend/infra engineering experience with distributed systems, ML infrastructure, and a strong product ownership mindset.
190k – 265k/yr
On-site8+ YOEML Engineering
Staff Software Engineer, AI Runtime
DatabricksMountain View, CA +1
Staff Software Engineer building and scaling Databricks' managed large-scale GPU training platform (AIR). Focus on distributed training performance, scheduling, fault tolerance, and developer experience for thousands of accelerators.
190k – 265k/yr
On-site10+ YOEML Engineering
Staff Machine Learning Engineer
DatabricksSan Francisco, CA +1
Develops and deploys state-of-the-art GenAI models and systems for Databricks products like Assistant and Genie. Requires 2-8 years ML engineering experience, proficiency in Python/PyTorch/TensorFlow, and expertise in LLMs.