
hud
San Francisco, CA
Platform for building RL environments and agent evals
About
HUD builds reinforcement learning environments and evaluation platforms for AI agents, turning software, web apps, and interfaces into testable setups. It serves frontier AI labs, researchers, and enterprises to benchmark, train, and improve agent performance at scale. This infrastructure enables reliable real-world AI agent deployment through detailed telemetry and scalable evals.
Perks & benefits
Health insurance, Free lunch, Gym membership, Commuter benefits, 401k
More AI companies
AI companiesOpen jobs
13Own and build the company’s security program as its first full-time security hire, covering product, cloud, infrastructure, incident response, compliance, and customer trust. The role requires hands-on security engineering and incident leadership, with experience operating SOC 2 or comparable frameworks.
Own and scale HUD’s legal operations infrastructure while drafting and negotiating commercial technology agreements. The role combines contracting judgment with workflow automation, legal technology, and strategic outside-counsel management in a fast-growing AI startup.
Build end-to-end product surfaces, backend services, internal tools, vendor workflows, dashboards, and observability systems powering an RL data engine. The role requires strong full-stack skills, especially Python and modern web development, plus comfort with cloud infrastructure and ambiguous product requirements.
Lead the data quality team at HUDHUD to build QC systems, validation methods, and experiments that measure and improve training data for frontier AI agents and RL environments. Requires deep data quality intuition, Python/Docker/Linux proficiency, and experience turning research insights into production evaluation pipelines.
Build synthetic data pipelines and tasks to train frontier AI agents. Requires Python, Docker, Linux, and experience creating realistic, scalable synthetic training data for models and evals.
Build high-quality, domain-specific benchmarks and infrastructure to rigorously evaluate frontier AI agents on realistic workflows. Requires strong Python/Docker/Linux skills, experience with evals or benchmarks, and a deep understanding of what makes a benchmark reliable and useful.
Build QC automation systems for RL training data and agent evals at HUDHUD. Design quality standards, validation pipelines, experiments and metrics without heavy LLM reliance; partner with vendors to debug and improve data generation. Requires Python, Docker, Linux and experience building scalable QA/QC systems end-to-end.
Applied Research Engineer owning technical deployment requests, troubleshooting, building tools/pipelines, and resolving ambiguous problems for frontier AI labs and data vendors at HUDHUD. Requires strong Python/Docker/Linux skills, eval/benchmark experience, independent problem-solving in fast-paced ambiguous settings.
Recruiter to own end-to-end hiring for engineering, product, and GTM roles at an early-stage AI startup. Requires experience sourcing and evaluating technical talent, building recruiting processes and tooling, and strong product understanding to pitch candidates.
Build and own end-to-end outbound GTM pipeline for AI training data platform, including technical outreach to ML engineers and AI researchers, workflow automation with AI tools, and deal closing. Requires engineering background, GTM tooling expertise, and startup experience.
Head of Growth at HUD to evangelize frontier AI infrastructure vision by building networks with AI researchers/engineers, leading recruiting and culture definition, hosting events, and creating content. Ideal for extraverted networkers with startup experience who can learn recruiting, marketing, and GTM functions.
Owns marketing strategy, social media growth on Twitter/X and LinkedIn, content creation explaining technical AI products, and event planning/execution to build brand and pipeline for AI infrastructure startup. Requires startup marketing experience and technical communication skills.
Builds QA systems, tooling, and workflows to audit and validate large-scale RL training data from suppliers. Partners with vendors to improve data quality using Python, Docker, and AI/ML techniques for frontier AI infrastructure.