Senior System Engineer
Build scalable infrastructure, automation, and network-resilience tools for a global cloud network. The role requires 6+ years in SRE, DevOps, or software engineering, strong Linux and distributed-systems expertise, programming skills in Python or Go, and experience with observability and IaC.
About the job
Responsibilities
- Build and maintain software tools, infrastructure, and services that improve network resilience and reduce operational toil.
- Develop practical, maintainable, scalable, and efficient solutions across new and existing systems.
- Leverage AI tools and automated assistants to streamline workflows and support project delivery.
- Design reliable and fault-tolerant distributed systems.
Requirements
- Bachelor’s degree in Computer Science or equivalent experience.
- 6+ years of overall experience in SRE, DevOps, or software engineering.
- Deep understanding of modern Linux internals and distributed systems infrastructure.
- Experience with containerization using Docker and Kubernetes.
- Experience with infrastructure as code tools and CI/CD pipeline automation.
- Proficiency in Python or Go.
- Hands-on experience with Prometheus, Grafana, distributed tracing, and metrics-driven alerting.
- Experience working with AI assistants or automated tools to augment task execution and project delivery.
Nice-to-haves
- Networking engineering knowledge, including Layer 2 and Layer 3 protocols and network APIs such as gRPC/gNMI and NETCONF.
- Experience with Ansible, Chef, or SaltStack.
- Experience with declarative configuration management and state-reconciliation systems.
- Experience managing internal or external customer requirements and expectations.
Compensation
- Estimated annual salary: €66,000–€91,000 for Portugal-based hires.
- Eligible to participate in the company’s equity plan.
Skills
Linux, Docker, Kubernetes, Infrastructure As Code, CI/CD, Python, Go, Prometheus, Grafana, Distributed Tracing, gRPC, Netconf, Ansible, Chef, Saltstack
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.