Senior DevOps Engineer
Owns and automates multi-cloud, multi-region SaaS infrastructure, focusing on reliability, observability, performance, and incident response. Requires 5–7+ years of DevOps or infrastructure engineering experience, cloud tooling expertise, and business-level English and Japanese.
About the job
Responsibilities
- Own the deployment, health, and continuous improvement of Tulip's multi-cloud, multi-region SaaS environments, including clusters spanning the US, Europe, and Asia.
- Design and evolve cloud architecture to ensure customer availability, stability, and performance as Tulip scales globally.
- Own and continuously improve CI/CD infrastructure, driving toward a fully automated, human-interaction-free software delivery lifecycle.
- Build automation tooling and internal systems that reduce operational toil and increase developer velocity.
- Define and maintain observability standards across cloud environments, including metrics, alerting, logging, and distributed tracing.
- Proactively identify performance degradation and capacity risks before they impact customers; lead incident response and drive root cause analysis.
- Partner closely with application engineering teams throughout the software development lifecycle, providing infrastructure guidance and support.
- Participate in the on-call rotation and contribute to continuous improvement through documentation, runbooks, and process iteration.
Requirements
- 5–7+ years of hands-on DevOps or infrastructure engineering experience, with demonstrated ownership of production cloud environments at scale.
- Proficiency with modern cloud infrastructure tooling, including Kubernetes, Helm, Terraform, Ansible, and major cloud providers such as AWS and/or Azure.
- Experience managing enterprise-grade data persistence layers, including NoSQL and SQL databases, key/value stores, and messaging systems such as AMQP and MQTT.
- Familiarity with observability and monitoring tooling such as Prometheus, Mimir, Thanos, and Grafana, with a strong understanding of SRE practices in a fast-growing SaaS environment.
- Exposure to programming or scripting languages used in infrastructure contexts, such as Go, TypeScript, Python, and Bash.
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- Business-level English for internal communication and English-language knowledge-base work.
- Business-level Japanese for communication with Japanese customers and internal teams, and for handling Japanese-language materials.
- Valid Japanese work visa at the time of application; visa sponsorship is not available.
Benefits and Compensation
- Company equity.
- Learning and Development benefit.
- Comprehensive social insurance package, including health, pension, employment, and workers' compensation insurance.
- Commutation allowance.
- Remote work environment in Japan for 2026, with future plans to open an office in 2027.
- The posting also references a hybrid work style.
Skills
Kubernetes, Helm, Terraform, Ansible, AWS, Azure, Prometheus, Mimir, Thanos, Grafana, Go, TypeScript, Python, Bash, CI/CD
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Designs and operates highly available GCP infrastructure and developer platforms for trading-critical systems. The role requires 5+ years of DevOps, platform, infrastructure, or SRE experience, with strong Terraform, Kubernetes, networking, CI/CD, observability, and incident-management skills.