Senior DevOps Engineer/SRE
Build and operate scalable infrastructure powering FlexAI’s AI and PaaS platform. The role focuses on Kubernetes, infrastructure as code, CI/CD, observability, incident response, and reliability practices, requiring 4+ years of DevOps, SRE, or infrastructure engineering experience.
About the job
Responsibilities
Infrastructure and Operations
- Build and maintain infrastructure for an AI and PaaS platform.
- Deploy and operate Kubernetes clusters and containerized services.
- Implement Infrastructure as Code using Pulumi or similar tools.
- Operate production systems at scale.
Reliability and SRE
- Define and implement SLIs, SLOs, and error budgets.
- Improve system reliability, availability, and performance.
- Participate in on-call rotations, incident response, and postmortems.
CI/CD and Automation
- Build and improve CI/CD pipelines for reliable, fast releases.
- Automate operational workflows and reduce manual toil.
- Contribute to GitOps and platform engineering practices.
Observability and Performance
- Implement and maintain observability using VictoriaMetrics and Grafana for metrics, logs, and traces.
- Monitor systems and troubleshoot latency, throughput, and cost issues.
Collaboration
- Work with developers, platform teams, and AI teams to support production systems.
- Debug issues across infrastructure and application layers.
- Improve engineering productivity and developer experience.
Requirements
- 4+ years of experience in DevOps, SRE, or infrastructure engineering.
- Experience operating production systems at scale.
- Hands-on experience with Kubernetes and containers.
- Experience with Infrastructure as Code, such as Pulumi or Terraform.
- Experience with cloud or hybrid environments, including AWS, Google Cloud, Azure, or on-premises infrastructure.
- Experience with observability tools such as Prometheus, Grafana, or OpenTelemetry.
- Experience with CI/CD systems and automation.
- Proficiency in Python, Go, or Bash.
- Strong debugging and problem-solving skills.
- Familiarity with SLOs and reliability practices.
- Experience working in startup or fast-paced environments.
- Comfort leveraging AI coding tools and agents.
Nice to Have
- Experience with AI/ML infrastructure or GPU workloads.
- Familiarity with distributed systems or compute platforms.
- Exposure to platform engineering concepts.
- Experience supporting systems from beta through production.
Benefits and Compensation
- Work on cutting-edge AI infrastructure.
- Build systems that power developers and enterprises.
- High ownership, fast execution, and real impact.
- Collaborative, high-caliber team.
Skills
Kubernetes, Containers, Pulumi, Terraform, AWS, GCP, Azure, Prometheus, Grafana, OpenTelemetry, CI/CD, GitOps, Python, Go, Bash
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.
Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.
Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.