# DevOps

**Company:** [Acryldata](https://hotfix.jobs/companies/acryldata)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** AWS, GCP, Microsoft Azure, Docker, Kubernetes, Terraform, CloudFormation, Python, Java, Prometheus, Grafana, Datadog, CI/CD, Infrastructure Automation, SaaS
**Posted:** 2026-09-03

> Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.

## Job Description

## Responsibilities

### Enterprise Platform Development
- Partner with product and engineering teams to influence advanced deployment capabilities.
- Build systems for seamless installation, upgrade, and rollback across varied environments.
- Design and implement monitoring and health-check systems for distributed deployments.
- Develop self-healing and automated remediation capabilities.

### Platform Reliability and Operations
- Establish and maintain SLAs and SLOs for cloud and enterprise offerings.
- Lead incident response and postmortem processes.
- Optimize system performance, capacity planning, and cost efficiency.
- Collaborate with product, engineering, and customer success teams to ensure reliable product delivery.
- Improve on-call practices, runbooks, and knowledge-sharing processes.
- Drive cross-functional initiatives to improve system reliability.

## Requirements

- 5+ years of experience in site reliability engineering, platform engineering, or DevOps.
- Strong expertise with cloud platforms such as AWS, Google Cloud, or Azure.
- Experience with infrastructure automation tools.
- Proficiency with Docker, Kubernetes, and container orchestration.
- Experience with infrastructure as code tools such as Terraform and CloudFormation.
- Strong programming skills in Python, Java, or similar languages.
- Familiarity with monitoring and observability tools such as Prometheus, Grafana, or Datadog.
- Experience with CI/CD pipelines and deployment automation.

## Nice-to-Haves

- Experience building and operating multi-tenant SaaS platforms.
- Background developing customer-facing deployment and management tools.
- Knowledge of data infrastructure and metadata management systems.

## Compensation and Benefits

- Competitive compensation.
- Equity for every team member.
- Remote-work support and a monthly coworking stipend.
- Comprehensive medical, dental, and vision coverage.
- Flexible spending accounts, including dependent care options.
- Fertility and family-forming support for U.S. employees.
- Unlimited paid time off and sick leave.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Site Reliability Engineer](https://hotfix.jobs/jobs/7db1f7e8-8d55-478b-9857-eb3bb0901fa0) - Invisible Tech - Remote
- [Platform Engineer - Compute Capacity](https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4) - Supabase - Remote
- [Production Support Engineer](https://hotfix.jobs/jobs/328c6ec8-1e18-42d6-99c0-515d7801c9ac) - Alpaca - Remote
- [Software Engineer - Platforms & Productivity](https://hotfix.jobs/jobs/39485182-8244-4096-81e1-8cc0ca6e1277) - Cloudflare

**Apply:** https://hotfix.jobs/jobs/d1d4b180-df1a-4699-b783-d501b31510b4
**Canonical:** https://hotfix.jobs/jobs/d1d4b180-df1a-4699-b783-d501b31510b4