Skip to content
AirbnbAirbnbUnited States

Operations Engineer, BizTech

Operations Engineer using AI, LLMs, and intelligent automation for ticket triage, incident response, self-healing observability, and workflow automation in Airbnb's BizTech Global Operations team. Requires 3+ years in observability, data pipelines, IaC, and CI/CD across clouds.

136k – 160k/yr
Remote3+ YOEDevOps / SRE

About the role

Responsibilities

  • Manage the ticket queue, prioritizing and resolving requests while identifying recurring categories to automate or deflect.
  • Participate in a rotating on-call and incident response schedule, including weekends, troubleshooting and documenting issues in real time.
  • Build and maintain monitoring dashboards (Tableau, Superset, Grafana) that track service health, availability, and data quality, flagging gaps as they emerge.
  • Use AI-assisted tools (e.g., Claude, Copilot) to speed up triage, root-cause analysis, scripting, and documentation — while knowing when a problem needs hands-on judgment instead.
  • Write and maintain code in general-purpose languages (Python, Go, JavaScript, TypeScript, Bash) to automate operational workflows.
  • Partner with global stakeholder teams to drive issues to resolution.
  • Apply AI at the forefront of operational health: using LLM-powered triage and intelligent automation to resolve tickets, speed up incident response, and build self-healing observability.
  • Prototype agentic workflows, embed AI into runbooks and diagnostics, and continuously look for repetitive work automation can take over.

Requirements

  • 3+ years of experience with observability and metrics tooling (e.g., Prometheus, Grafana, Datadog, ElasticSearch).
  • 3+ years working with data querying and pipelines (e.g., SQL, Airflow, Trino, SQS).
  • Working knowledge of network fundamentals and hardware (e.g., Cisco, Palo Alto).
  • Experience with Infrastructure as Code.
  • Hands-on experience with CI/CD and automation tooling (e.g., Jenkins, ArgoCD, GitHub Actions) across AWS, GCP, or Oracle Cloud.
  • Comfort working in ticket/workflow-driven environments using Jira and Confluence.

Nice-to-Haves

  • Experience supporting enterprise SaaS tools (e.g., Salesforce, Workday) and SSO/identity integration (e.g., Okta).

Skills

PrometheusGrafanaDatadogElasticsearchSQLAirflowtrinoSQSInfrastructure As CodeJenkinsArgo CDGitHub ActionsAWSGCPJira

Similar roles

DevOps / SRE jobs
Upstart

DevOps Engineer, Cloud Platform

UpstartUnited States

Build and operate shared Kubernetes (EKS) and AWS cloud infrastructure powering Upstart's product and ML workloads. Requires 3+ years Kubernetes production experience plus strong AWS, IaC, and GitOps skills.

136k – 197k/yr
Remote5+ YOEDevOps / SRE
Baseten

Forward Deployed SRE

BasetenSan Francisco, CA +1

Site Reliability Engineer owns reliability of multi-cloud Kubernetes infrastructure for AI/ML platform, builds observability tooling as code, automates mitigations, leads incident response, and defines SLOs/SLIs. Requires extensive Kubernetes and observability experience.

135k – 285k/yr
HybridDevOps / SRE
Assembled

Software Engineer - Platform

AssembledNew York, NY

Build scalable infrastructure, integrations, and data platforms powering workforce management and AI agent products at enterprise scale. Requires 5+ years in backend/platform systems, with expertise in AWS, Kubernetes, Go/Python, and datastores like Postgres and Snowflake.

135k – 280k/yr
On-site5+ YOEDevOps / SRE
Ema

Software Engineer, DevOps

EmaPalo Alto, CA +1

Designs and builds scalable infrastructure for AI products, focusing on cloud platforms, Kubernetes orchestration, CI/CD pipelines, and observability. Requires 3+ years in infrastructure engineering and bachelor's/master's in CS.

135k – 225k/yr
Hybrid3+ YOEDevOps / SRE
Alchemy

Cloud Infrastructure Engineer

AlchemySan Francisco, CA +1

Designs, deploys, and improves scalable blockchain infrastructure using Kubernetes, Terraform, and cloud tools. Drives AI enablement, builds observability with Prometheus/Grafana, manages multi-cloud networks, and leads incident response. Requires 5+ years in SRE/infrastructure with strong automation focus.

135k – 240k/yr
Hybrid5+ YOEDevOps / SRE