Senior DevOps Engineer II
This senior DevOps role builds and operates AWS infrastructure for a SaaS platform and agentic AI services, with responsibility for reliability, observability, security, compliance, and disaster recovery. It requires 6+ years of platform experience, AI/ML infrastructure expertise, Terraform proficiency, and technical leadership.
About the job
Responsibilities
- Design and operate AWS cloud infrastructure for the core SaaS platform and agentic AI services, focusing on reliability, scalability, and cost efficiency.
- Build and maintain AI/ML infrastructure and monitoring for LLM-powered agentic services.
- Establish infrastructure-as-code standards with Terraform, including environment parity, drift detection, and automated compliance validation.
- Implement observability for data integrity, SLOs, error budgets, and automated regression detection.
- Build deployment automation with pre-deployment verification, migration validation, and automated rollback procedures.
- Support data pipelines, Redshift warehousing, analytics, BI, and AI training workflows.
- Implement security and compliance controls for AI workloads, including audit logging, access governance, and configuration management.
- Improve disaster recovery through documented procedures, defined RTO/RPO targets, and tested recovery runbooks.
- Lead architecture reviews for services, integrations, and AI agent deployments in partnership with engineering, product, and security.
- Improve developer experience across testing environments, CI/CD pipelines, and local development workflows.
- Provide technical leadership for infrastructure decisions and mentor engineers on operational practices.
Requirements
- 6+ years of DevOps, SRE, or platform engineering experience.
- 2+ years building or operating AI/ML infrastructure, including model serving, inference, LLM orchestration, or agentic systems.
- Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
- Experience operating infrastructure for traditional and AI/ML workloads at a SaaS company.
- Deep AWS experience, including services such as ECS, Lambda, SageMaker, Bedrock, S3, DynamoDB, or Redshift.
- Strong Terraform skills, including state, modules, and multi-environment configuration management.
- Understanding of data pipelines, warehousing, ETL/ELT, and analytics at scale.
- Experience with data integrity monitoring, SLOs, error budgets, and silent-failure detection.
- Experience in compliance-sensitive environments and knowledge of audit trails, access governance, and change management.
- Strong proficiency in Python, Go, TypeScript, or a similar programming language.
- Ability to lead cross-team technical initiatives and mentor engineers.
- Strong communication skills with technical and non-technical stakeholders.
Nice to Have
- Experience with ECS or EKS.
- Experience with Datadog or CloudWatch.
- Experience with Redshift or similar data warehouses.
- Experience with Airflow, dbt, Dagster, or similar pipeline tools.
- Experience in insurance, financial services, healthcare, or another regulated industry.
- Curiosity about AI and emerging technologies, with sound judgment in applying them responsibly.
Compensation & Benefits
- Annual base salary: $119,000–$221,000 USD.
- Equity grants for all employees.
- 4% matching 401(k) program.
- Medical, dental, vision, disability, and life insurance for employees working 30+ hours per week.
- Monthly wellness stipend.
- Paid parental leave.
- Flexible vacation policy.
Skills
AWS, Terraform, Amazon Ecs, Amazon Eks, AWS Lambda, Amazon Sagemaker, Amazon Bedrock, Amazon S3, Amazon Dynamodb, Amazon Redshift, Datadog, Amazon Cloudwatch, Airflow, dbt, Dagster
Similar jobs
DevOps / SRE jobsSenior network engineer responsible for designing, operating, and securing MongoDB’s global network and VPN infrastructure. The role requires 6+ years of networking or systems engineering experience, strong enterprise networking expertise, automation skills, and the ability to lead complex infrastructure initiatives.
Senior Site Reliability Engineer responsible for designing and operating reliable, scalable production infrastructure, leading incident response, and improving observability and resilience. Requires 5+ years of reliability-focused engineering experience and expertise across cloud, infrastructure as code, Kubernetes, monitoring, and application development.
Leads end-to-end infrastructure for a scientific imaging platform, covering Linux administration, GPU/HPC systems, storage, upgrades, and vendor coordination. The role supports AI-enabled imaging workflows and requires extensive production Linux, Image Artist, GPU, HPC, and enterprise storage experience.
Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.