Software Engineer, Infrastructure
The engineer will evolve and operate Notion’s async task runner and configuration management platform, supporting reliability and scalability for more than 100 million users. The role requires at least four years of software development experience and knowledge of distributed systems, production operations, and infrastructure tradeoffs.
About the job
Responsibilities
- Contribute to the evolution and maintenance of Notion’s internal async task runner, supporting more than 100 million global users.
- Maintain and scale the configuration management platform for product and infrastructure engineers.
- Evaluate and integrate technologies to address emerging infrastructure challenges.
- Debug live production systems with minimal disruption, including replacing components and managing failovers.
- Participate in an on-call rotation, respond to incidents, and restore normal operations quickly.
- Partner with product engineers and infrastructure teams to align platforms with business and product needs.
Requirements
- At least 4 years of software development experience.
- Understanding of distributed systems and tradeoffs involving consistency, latency, and scalability.
- Pragmatic problem-solving skills, with attention to business impact, maintainability, and delivery speed.
- Ability to take ownership of ambiguous problems and work effectively in a fast-paced, unstructured environment.
- Strong cross-functional collaboration and customer empathy.
- Curiosity about and willingness to use AI tools to improve productivity.
Benefits and Compensation
- Opportunity to work on infrastructure supporting more than 100 million global users.
- Participation in an on-call rotation.
- Opportunity to develop expertise in distributed systems and infrastructure.
Skills
Distributed Systems, Configuration Management, Async Task Runners, Infrastructure, Production Debugging, Failover Management, Scalability, Consistency, Latency, On-Call Operations
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.
Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.