Senior DevOps Engineer responsible for building and optimizing data streaming, processing, and monitoring infrastructure on AWS/GCP Kubernetes platforms. Requires 7+ years DevOps/SRE experience, Terraform, CI/CD tools, Python, and strong cloud-native data services knowledge.
133k – 209k/yr
Remote7+ YOEDevOps / SRE
About the role
What You'll Do
Drive initiatives to implement and enforce best practices for data streaming, processing, analytics and monitoring infrastructure.
Deploy and manage services on Kubernetes-based platforms such as Amazon EKS and Google Kubernetes Engine (GKE).
Provision and manage cloud infrastructure using Terraform, ensuring best practices in security, scalability, and cost-efficiency.
Maintain and optimize CI/CD pipelines using Jenkins, ArgoCD, and GitHub Enterprise Actions to support automated deployments and testing.
Work with cloud-native data services such as AWS Kinesis, AWS Glue, Google Dataflow, and Google Pub/Sub, BigQuery, BigTable.
Develop and maintain automation scripts and tooling using Python to support DevOps processes.
Monitor system performance, troubleshoot issues, and implement proactive solutions to enhance reliability and efficiency.
Implement SRE practices to improve service reliability, scalability, and cost-effectiveness.
Analyze and optimize cloud costs, identifying areas for improvement and implementing cost-saving strategies.
Ensure compliance with security policies and best practices in cloud environments.
Drive adoption of company standards and influence data teams to align with best DevOps and SRE practices.
Collaborate with cross-functional teams to improve development workflows and infrastructure.
Required Qualifications
7+ years of experience in a DevOps, Site Reliability Engineering, or Cloud Infrastructure role.
Strong experience with AWS and GCP data services, including Kinesis, Glue, Pub/Sub, and Dataflow.
Proficiency in deploying and managing workloads on Kubernetes (EKS/GKE) in production environments.
Hands-on experience with Infrastructure-as-Code (IaC) using Terraform.
Expertise in CI/CD pipeline management using Jenkins, ArgoCD, and GitHub Enterprise Actions.
Programming skills in Python for automation and scripting.
Experience with observability and monitoring tools (e.g., Prometheus, Grafana, Datadog, or CloudWatch).
Strong understanding of SRE principles, including performance monitoring, incident response, and reliability engineering.
Experience with cost optimization strategies for cloud infrastructure.
Self-motivated and driven, with a strong ability to influence and drive changes across multiple teams.
Ability to work collaboratively in an agile environment and support multiple teams.
Preferred Qualifications
Experience with data lake architectures and big data processing frameworks (e.g., Apache Spark, Flink, Snowflake, BigQuery).
Familiarity with event-driven architectures and message queues (e.g., Kafka, RabbitMQ).
Experience with workflow orchestration tools such as Apache Airflow and Google Cloud Composer.
Knowledge of service mesh technologies like Istio.
Experience with GitOps workflows and Kubernetes-native tooling.
Senior infrastructure engineer building scalable abstractions, platform tooling, and the full infrastructure stack (AWS to product platform) that accelerates Pilot's R&D and business growth. Requires 5+ years software engineering experience, production Python, Terraform/AWS, frontend familiarity, mentoring ability, and strong collaboration/communication skills.
133k – 267k/yr
Remote5+ YOEDevOps / SRE
Senior Network Engineer
Northwood SpaceLos Angeles, CA +1
Design, deploy, and operate enterprise network infrastructure for corporate facilities and hybrid cloud environments with zero-trust architecture and compliance requirements. Requires 5+ years enterprise networking experience and ability to obtain TS/SCI clearance.
133k – 215k/yr
On-site5+ YOEDevOps / SRE
Senior Cloud Software Engineer - AutoScaling
ClickhouseUnited States
Develops and maintains auto-scaling infrastructure including Kubernetes operators for ClickHouse cloud platform. Requires 5+ years building scalable distributed systems, expertise in Go/C++/Java, public clouds (AWS/GCP/Azure), and data tools like Spark/Kafka.
133k – 232k/yr
Remote5+ YOEDevOps / SRE
Senior AI-Enabled DevOps Engineer
PointClickCareUnited States
Designs, builds, and operates scalable cloud infrastructure and developer platforms supporting AI/ML workloads. Owns end-to-end systems, builds CI/CD pipelines, ensures reliability with observability tools, and integrates AI-assisted DevOps practices. Requires 5+ years experience with cloud platforms, Kubernetes, and IaC.
Builds and maintains automated testing infrastructure, CI/CD pipelines, and developer tools using Cypress and GitHub Actions/CircleCI. Mentors QA engineers and drives testing best practices in a legal SaaS platform. Requires 5+ years Cypress experience and expertise in major languages.