DevOps
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.
About the job
Responsibilities
Enterprise Platform Development
- Partner with product and engineering teams to influence advanced deployment capabilities.
- Build systems for seamless installation, upgrade, and rollback across varied environments.
- Design and implement monitoring and health-check systems for distributed deployments.
- Develop self-healing and automated remediation capabilities.
Platform Reliability and Operations
- Establish and maintain SLAs and SLOs for cloud and enterprise offerings.
- Lead incident response and postmortem processes.
- Optimize system performance, capacity planning, and cost efficiency.
- Collaborate with product, engineering, and customer success teams to ensure reliable product delivery.
- Improve on-call practices, runbooks, and knowledge-sharing processes.
- Drive cross-functional initiatives to improve system reliability.
Requirements
- 5+ years of experience in site reliability engineering, platform engineering, or DevOps.
- Strong expertise with cloud platforms such as AWS, Google Cloud, or Azure.
- Experience with infrastructure automation tools.
- Proficiency with Docker, Kubernetes, and container orchestration.
- Experience with infrastructure as code tools such as Terraform and CloudFormation.
- Strong programming skills in Python, Java, or similar languages.
- Familiarity with monitoring and observability tools such as Prometheus, Grafana, or Datadog.
- Experience with CI/CD pipelines and deployment automation.
Nice-to-Haves
- Experience building and operating multi-tenant SaaS platforms.
- Background developing customer-facing deployment and management tools.
- Knowledge of data infrastructure and metadata management systems.
Compensation and Benefits
- Competitive compensation.
- Equity for every team member.
- Remote-work support and a monthly coworking stipend.
- Comprehensive medical, dental, and vision coverage.
- Flexible spending accounts, including dependent care options.
- Fertility and family-forming support for U.S. employees.
- Unlimited paid time off and sick leave.
Skills
AWS, GCP, Microsoft Azure, Docker, Kubernetes, Terraform, CloudFormation, Python, Java, Prometheus, Grafana, Datadog, CI/CD, Infrastructure Automation, SaaS
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Build and operate internal developer platforms that improve engineering velocity, reliability, and security. The role spans developer tooling, CI/CD, GitOps, AI-assisted development, and automated engineering guardrails.