Software Engineer, Data Infrastructure
Architect and build foundational data infrastructure for massive simulation outputs. Design novel data models and high-throughput pipelines to feed LLMs with structured context from complex, state-based environments.
About the job
Key Responsibilities
- Architect the Data Ensemble: Design and implement the architecture to ensemble various sources of injected context (deeply structural simulation data, historical game states, and dynamic user inputs) into a unified, highly queryable format optimized for LLM consumption.
- Massive Batch Infrastructure: Build highly scalable, resilient data architectures from scratch. Optimize for moving, transforming, and processing massive quantities of simulation output data via enormous batch jobs, maintaining the minimal latency required for rapid wargame iterations.
- Complex Data Modeling: Design sophisticated, highly relational data models that accurately represent massive, state-based simulation environments, making them easily interpretable by machine learning models.
- First-Principles Problem Solving: Navigate highly ambiguous product requirements to design custom, ground-up systems where existing open-source or enterprise tools simply cannot handle the structural complexity or scale.
- Technical Leadership: Set the technical standard for the data infrastructure team, driving rigorous code quality, system performance, and architectural clarity.
Requirements
- 5+ years of backend or data infrastructure experience, operating at a Senior, Staff, or Principal level.
- Deep, expert-level proficiency in systems languages (e.g., Rust, Go, C++, or highly optimized Python/Java, Spark) and a fundamental understanding of memory management, compute limits, and distributed systems architecture.
- Proven track record of processing massive datasets. Understand how to optimize massive batch jobs and parallel processing across distributed simulation nodes without sacrificing speed.
- Expert in surfacing the right needle from an ocean of hay to feed decision-making engines. Backgrounds in Search & RecSys, Gaming / MMOs, or High-Frequency Trading (HFT) highly valued.
- Strong desire to build robust, foundational technology that supports national security and defense modernization.
Nice to Have
- Active Secret or TS/SCI clearance, or eligibility and willingness to obtain one.
- Experience with LLM context optimization, vector embeddings, or agentic AI frameworks (e.g., advanced RAG architectures).
- Deep domain experience working with wargaming data, complex systems modeling, or distributed simulation protocols.
- Previous experience in a high-growth, 0-to-1 startup environment.
Benefits
- Comprehensive health, dental and vision coverage
- Retirement benefits
- Learning and development stipend
- Generous PTO
- Commuter stipend (role-dependent)
Skills
Rust, Go, C++, Python, Java, Spark, Distributed Systems, Data Modeling, Batch Processing, Information Retrieval
Similar jobs
Data Engineering jobsBuilds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.
Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.