Principal / Staff / Senior Infrastructure Engineer
Own and scale secure cloud infrastructure, deployments, observability, compliance, and incident response for a hardware collaboration platform. The role requires substantial cloud or security engineering experience, AWS and Linux expertise, and the ability to lead cross-functional infrastructure initiatives.
About the job
Responsibilities
- Own server deployments and infrastructure automation.
- Manage enterprise deployments and SSO integrations with customer success.
- Partner with application developers on infrastructure solutions.
- Document security processes and identify improvements.
- Manage ISO 27001 and SOC 2 certification processes and penetration testing.
- Improve architecture and hosting services.
- Automate software deployments, backups, and data recovery.
- Monitor and improve cloud infrastructure performance and availability.
- Participate in on-call rotations, respond to incidents, and lead blameless postmortems.
- Scale infrastructure for zero-downtime deployments.
- Reduce cloud costs and improve performance.
- Establish disaster recovery processes and stress-test backups.
- Support CI/CD workflows.
- Identify security flaws and conduct penetration testing and red-team exercises.
Requirements
- 4+ years of cloud infrastructure and/or security engineering experience for Senior; 6+ years for Staff; 8+ years for Principal.
- Deep hands-on AWS and Linux administration experience.
- Experience with SOC 2, ISO 27001, and penetration testing.
- Strong project management and cross-functional leadership skills.
- Comfort with ambiguity and autonomy.
- Fluency with day-to-day AI tools.
- Bachelor's degree or higher in a technology-related field.
- Must be a U.S. Citizen or lawful permanent resident.
Nice-to-haves
- AWS security services including IAM, GuardDuty, and AWS WAF.
- Terraform, Ansible, and infrastructure-as-code.
- Logging aggregation and metrics observability.
- Docker, Kubernetes, and Helm.
- Distributed systems, secure multi-tenant SaaS, and self-hosted deployments.
- SBOM generation and auditing.
- Bash and Python scripting; Go or Rust programming.
- Nginx and reverse-proxy services.
- Documentation and project management.
- SSO, OIDC, and LDAP.
- Data lakes and high-availability services.
Compensation and Benefits
- Competitive salary and equity.
- Health, dental, and vision benefits.
- Generous paid time off.
- Home office stipend.
- Relocation package.
- Flexible work and flex-office availability.
Skills
AWS, Linux, Terraform, Ansible, Docker, Kubernetes, Helm, IAM, Guardduty, Aws Waf, Prometheus, Grafana, Python, Bash, Postgres
Similar jobs
DevOps / SRE jobsThis principal-level role owns operational excellence for a hyperscale AI data center network fleet, leading readiness, high-risk changes, audits, and incident resolution across sites. It requires extensive mission-critical network operations experience, routing and optical networking expertise, and 50–75% travel.
Leads Snowflake’s cloud infrastructure performance strategy by evaluating new hardware, building benchmark and validation systems, and translating performance data into pricing, capacity, and rollout decisions. Requires 12+ years in performance, systems, or infrastructure engineering and deep cloud hardware expertise.
As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.
Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.