Staff+ Software Engineer, Safeguards Data
Build and operate the data foundations behind AI safeguards, including production pipelines, data stores, governance controls, and internal tooling. The role requires strong Python and SQL skills, production data-platform experience, and expertise in reliable, privacy-conscious systems across multiple cloud environments.
About the job
Responsibilities
- Build and operate the Safeguards data platform, including ingestion and processing pipelines, warehouses and other data stores, and the schemas and interfaces used by detection and review systems.
- Keep Safeguards systems running day to day and maintain a high operational bar for safety and customers while reducing manual effort.
- Own data governance and integrity, including retention and access controls, privacy-preserving handling of sensitive data, lineage, and correctness guarantees.
- Design systems that run portably across cloud providers and within customer-managed and third-party environments.
- Partner with analysts, investigators, and researchers to build the internal tooling they depend on.
Minimum Qualifications
- Proficiency in Python and SQL.
- Experience building and operating data pipelines or data stores in production.
- Ability to work across the data stack, including ingestion, storage, and consumption.
- Strong written and verbal communication skills, including the ability to explain technical tradeoffs to people outside your discipline.
Preferred Qualifications
- Extensive software engineering experience, including significant work on data-intensive systems.
- Experience with integrity, spam, fraud, or abuse detection and mitigation.
- Experience building trust and safety detection and intervention mechanisms for AI or machine learning systems.
- Experience building and operating large-scale distributed data infrastructure.
- Experience working across multiple cloud providers or building provider-agnostic infrastructure.
- Experience meeting data governance requirements in a regulated or high-sensitivity domain.
- Experience working closely with operational teams to build custom internal tooling.
Compensation and Benefits
- Annual salary: $320,000–$485,000 USD.
- Minimum education: bachelor’s degree or an equivalent combination of education, training, and/or experience.
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Office collaboration space.
- Visa sponsorship may be available for eligible roles and candidates.
Skills
Python, SQL, Data Pipelines, Data Stores, Data Warehousing, Data Governance, Access Control, Data Lineage, Distributed Systems, AWS, GCP, Microsoft Azure, Fraud Detection, Abuse Detection, Internal Tooling
Similar jobs
Data Engineering jobsLeads the design and delivery of scalable data infrastructure, services, and developer tooling while shaping technical direction and mentoring engineers. Requires 8+ years of software engineering experience, strong systems design expertise, and proficiency in a modern programming language.
Staff Data Platform Engineer leading the architecture and development of financial data infrastructure for revenue reporting, billing, forecasting, and compliance. Requires 8+ years of data engineering or architecture experience, strong streaming and warehouse expertise, and the ability to mentor engineers and partner with Finance and Audit leaders.
Leads the re-platforming of Vanta’s compliance data layer from MongoDB to schema-aware PostgreSQL across high-throughput Kafka and S3 pipelines. The role requires staff-level distributed systems expertise, migration leadership, and strong experience with relational and document data modeling.
Staff Software Engineer responsible for designing and scaling Ray Data’s distributed data-processing infrastructure for large-scale AI training and inference. Requires 6+ years of production software and architectural ownership experience, plus deep distributed-systems expertise and strong Python skills.
Staff-level engineer leading backend services and data-platform architecture, including large-scale ingestion, distributed systems, and trustworthy BigQuery/dbt warehouse models. Requires 10+ years of software engineering experience, expert Python, deep SQL/dbt expertise, and strong technical leadership.