Senior Site Reliability Engineer
Senior Site Reliability Engineer automates operational processes, manages secure infrastructure across hybrid data centers and cloud, and improves workflows for ML/data teams. Requires 3-5 years experience with Linux, containers, and automation tools.
About the job
Responsibilities
- Automate manual operational processes
- Improve workflows of developer, data, and machine learning teams
- Manage secure integration and deployment tooling
- Create, maintain, monitor, and audit secure infrastructure
- Manage a diverse array of technology platforms, following best practices and procedures
- Participate in on-call rotation and root cause analysis
- Maintain awareness of industry best practices for data maintenance handling as it relates to your role
- Adhere to policies, guidelines and procedures pertaining to the protection of information assets
- Report actual or suspected security and/or policy violations/breaches to an appropriate authority
Requirements
- Minimum 3 - 5 years of previous experience in development, operations, IT, or a related field
- Comfortable working on Linux infrastructures (Debian) via the CLI
- Able to learn quickly in a fast-paced environment
- Able to debug, optimize, and automate routine tasks
- Able to multitask, prioritize, and manage time efficiently independently
- Able to physically lift equipment at least 30 pounds
- Can communicate effectively across teams and management levels
- Degree in computer science, or similar, is an added plus
Technology Stack
Operating Systems: Linux/Debian Family/Ubuntu
Configuration Management: Chef
Containerization: Docker
Container Orchestrators: Mesosphere/Kubernetes
Scripting Languages: Python/Ruby/Node/Bash
CI/CD Tools: Jenkins
Network hardware: Arista/Cisco/Fortinet
Hardware: HP/SuperMicro
Storage: Ceph, S3
Database: Scylla, Postgres, Pivotal GreenPlum
Message Brokers: RabbitMQ
Logging/Search: ELK Stack
AWS: VPC/EC2/IAM/S3
Networking: TCP/IP, ICMP, SSH, DNS, HTTP, SSL/TLS, Storage systems, RAID, distributed file systems, NFS/iSCSI/CIFS
Skills
Kubernetes, Docker, Linux, Python, AWS, Chef, Jenkins, Elk Stack, Ceph, RabbitMQ
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.