Senior Software Engineer, Enterprise Resilience
Build and operate resilient systems for Vanta's FedRAMP and enterprise environments, define reliability frameworks, and partner with teams to ensure scalable, compliant infrastructure using AWS and modern tooling.
About the job
Responsibilities
- Build and operate systems powering Vanta’s FedRAMP environments, including automated release, vulnerability remediation, and evidence generation pipelines.
- Design and maintain vulnerability management platform for detection, remediation, and compliance reporting.
- Define production reliability framework including SLOs, incident response, observability standards, service catalog, metrics dashboards, and SLA.
- Improve incident response workflows and CI/deploy processes.
- Collaborate with product teams on reliability best practices and operational readiness.
- Lead datacenter and environment build-outs for FedRAMP expansion.
- Solve scalability and performance challenges.
Requirements
- Experience operating services in strict compliance environments like FedRAMP.
- Technical lead on large-scale reliability initiatives across engineering organizations.
- Technical leadership on infrastructure/platform teams.
- Experience with AWS services and scaling platforms.
- Strong focus on resilient, scalable services and thoughtful trade-offs.
- Open to using AI responsibly.
Compensation & Benefits
- Industry-competitive salary and equity.
- Comprehensive medical, dental, vision coverage (100% employee premiums covered for most plans).
- 16 weeks paid parental leave.
- Health & wellness, remote workspace stipends.
- Matching 401(k), flexible PTO, 11 holidays.
Skills
FedRAMP, AWS, TypeScript, Node.js, MongoDB, GitHub Actions, Fargate, ECS, SLOs, Observability
Similar jobs
DevOps / SRE jobsBuild and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Leads design, deployment, and operation of secure distributed cloud systems for public-sector and air-gapped environments. Requires active or obtainable TS/SCI clearance with polygraph, U.S. citizenship, and 7+ years of production experience.