CloudOps Engineer
Design, build, and automate secure AWS cloud-native infrastructure with Kubernetes and Terraform. Enable dev teams with self-service platforms, CI/CD pipelines, and SRE best practices.
About the job
Principal Responsibilities
- Design, develop, integrate, test, and monitor all AWS infrastructure following industry best practices and standards in a secure and highly regulated environment
- Utilize AI-based tools for Infrastructure as Code Development
- Engineer self-service and automated solutions to enable application development teams to constantly move faster while ensuring compliance with security standards and requirements
- Build and pilot a next generation application platform centered around Kubernetes and its ecosystem of software
- Act as a trusted partner for application development teams, advising and supporting their use of cloud technologies
- Design, develop, and maintain CI/CD pipelines
- Apply secure by design and secure by default patterns to AWS and Kubernetes infrastructure and platforms
- Champion cloud-native best practices and partner with application development teams on their implementation
- Automate repetitive tasks and processes
Minimum Qualifications
- Bachelor's Degree or equivalent working experience
- 3+ years of professional experience
- Strong programming and scripting skills with one or more languages (Python, Go, Bash)
- Experience utilizing and deploying applications to a container orchestration platform like Kubernetes in a production environment
- Experience working with Linux operating systems and containers
Preferred Qualifications
- MS or PhD in Computer Science, Math, Physics or Engineering
- Professional experience designing and managing infrastructure in AWS with Terraform or other Infrastructure as Code tools in a highly regulated environment
- Experience implementing or integrating with application metric, logging, and tracing aggregation tools such as New Relic, Prometheus, OpenTelemetry, Loki, and Grafana for troubleshooting, monitoring, and alerting
- Experience using AI agents, such as GitHub Copilot, for development
- Practical working experience with event driven architectures and message buses like Kafka
- Experience owning CI/CD pipelines and automated workflows
- Ability to design and develop event driven automations and/or workflows using AWS Lambda and/or Step Functions
- Understanding of the SRE principles and/or a background in an SRE, operations, or system administrator role
Skills
AWS, Terraform, Kubernetes, Python, Go, Bash, Linux, CI/CD, Docker, Prometheus, Grafana, OpenTelemetry, Kafka, AWS Lambda, Step Functions
Similar jobs
DevOps / SRE jobsBuild and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.
Leads on-site deployment of data center physical infrastructure, managing contractors, performing QA/QC on fiber optics and cabling, and ensuring compliance with standards. Requires 5+ years experience, SME-level fiber optic expertise, bachelor's degree, and 40% travel readiness.
The Python Engineer will improve and operate trading systems, support integrations with asset classes and prime brokers, and handle monitoring, incidents, and performance optimization. The role requires 3+ years of experience, strong Python and Linux skills, and familiarity with market data and order-entry systems.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
The DevOps Engineer will build and operate reliable infrastructure, deployment workflows, and observability for data pipelines and AI/ML systems. The role requires at least three years of DevOps, SRE, or infrastructure experience plus strong cloud, Terraform, containerization, and MLOps expertise.