# Senior DevOps Engineer II

**Company:** [Hi Marley](https://hotfix.jobs/companies/hi-marley)
**Location:** Boston, MA
**Role:** DevOps / SRE
**Salary:** $119k – $221k/yr
**Experience:** 6+ years
**Skills:** AWS, Terraform, Amazon Ecs, Amazon Eks, AWS Lambda, Amazon Sagemaker, Amazon Bedrock, Amazon S3, Amazon Dynamodb, Amazon Redshift, Datadog, Amazon Cloudwatch, Airflow, dbt, Dagster
**Posted:** 2026-06-16

> This senior DevOps role builds and operates AWS infrastructure for a SaaS platform and agentic AI services, with responsibility for reliability, observability, security, compliance, and disaster recovery. It requires 6+ years of platform experience, AI/ML infrastructure expertise, Terraform proficiency, and technical leadership.

## Job Description

## Responsibilities
- Design and operate AWS cloud infrastructure for the core SaaS platform and agentic AI services, focusing on reliability, scalability, and cost efficiency.
- Build and maintain AI/ML infrastructure and monitoring for LLM-powered agentic services.
- Establish infrastructure-as-code standards with Terraform, including environment parity, drift detection, and automated compliance validation.
- Implement observability for data integrity, SLOs, error budgets, and automated regression detection.
- Build deployment automation with pre-deployment verification, migration validation, and automated rollback procedures.
- Support data pipelines, Redshift warehousing, analytics, BI, and AI training workflows.
- Implement security and compliance controls for AI workloads, including audit logging, access governance, and configuration management.
- Improve disaster recovery through documented procedures, defined RTO/RPO targets, and tested recovery runbooks.
- Lead architecture reviews for services, integrations, and AI agent deployments in partnership with engineering, product, and security.
- Improve developer experience across testing environments, CI/CD pipelines, and local development workflows.
- Provide technical leadership for infrastructure decisions and mentor engineers on operational practices.

## Requirements
- 6+ years of DevOps, SRE, or platform engineering experience.
- 2+ years building or operating AI/ML infrastructure, including model serving, inference, LLM orchestration, or agentic systems.
- Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
- Experience operating infrastructure for traditional and AI/ML workloads at a SaaS company.
- Deep AWS experience, including services such as ECS, Lambda, SageMaker, Bedrock, S3, DynamoDB, or Redshift.
- Strong Terraform skills, including state, modules, and multi-environment configuration management.
- Understanding of data pipelines, warehousing, ETL/ELT, and analytics at scale.
- Experience with data integrity monitoring, SLOs, error budgets, and silent-failure detection.
- Experience in compliance-sensitive environments and knowledge of audit trails, access governance, and change management.
- Strong proficiency in Python, Go, TypeScript, or a similar programming language.
- Ability to lead cross-team technical initiatives and mentor engineers.
- Strong communication skills with technical and non-technical stakeholders.

## Nice to Have
- Experience with ECS or EKS.
- Experience with Datadog or CloudWatch.
- Experience with Redshift or similar data warehouses.
- Experience with Airflow, dbt, Dagster, or similar pipeline tools.
- Experience in insurance, financial services, healthcare, or another regulated industry.
- Curiosity about AI and emerging technologies, with sound judgment in applying them responsibly.

## Compensation & Benefits
- Annual base salary: **$119,000–$221,000 USD**.
- Equity grants for all employees.
- 4% matching 401(k) program.
- Medical, dental, vision, disability, and life insurance for employees working 30+ hours per week.
- Monthly wellness stipend.
- Paid parental leave.
- Flexible vacation policy.

## Similar jobs

- [Senior Network Engineer](https://hotfix.jobs/jobs/b23e626b-5598-4d12-964a-efca91f98bfd) - MongoDB - Palo Alto, CA - $118k – $231k/yr
- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/a6949459-b2bc-44d0-99bf-23c9890dcf26) - PrizePicks - Remote - $120k – $175k/yr
- [Lead Scientific Imaging Systems Engineer](https://hotfix.jobs/jobs/c09c30f6-dd1b-4de5-aeb3-248a726a5186) - Axle - Rockville, MD - $120k – $145k/yr
- [Senior Software Engineer, Site Reliability](https://hotfix.jobs/jobs/0cb9abcb-6159-44b6-a3ae-bef58f9fdbdb) - Bloomerang - Remote - $115k – $150k/yr
- [Senior Site Infrastructure Engineer](https://hotfix.jobs/jobs/847213bd-91b8-44f9-b450-477e4aa3714e) - Shield AI - Seattle, WA - $110k – $210k/yr

**Apply:** https://hotfix.jobs/jobs/52b8cf0f-d5c0-43fa-b37a-fa2f2d8cb005
**Canonical:** https://hotfix.jobs/jobs/52b8cf0f-d5c0-43fa-b37a-fa2f2d8cb005