Build and maintain distributed systems, tooling, and platforms to measure service reliability, SLOs, and interactions at Cloudflare's massive scale. Requires strong programming in Go/Rust/Python and experience with highly-available systems, testing, and cross-team collaboration.
Salary not listed
HybridFullstack Engineering
About the role
Responsibilities
Build and run tooling and internal platform for engineering teams to measure service and feature reliability.
Develop and maintain distributed systems to verify interactions between systems and products in production at scale.
Communicate effectively with engineers across the company to understand system behaviors and deliver monitoring and reliability tooling.
Work closely with product managers and Product Site Reliability Engineers on quality of service measurements for enterprise customers.
Design, implement, and maintain secure and highly-available distributed systems.
Develop, document, and execute test and SLO plans, test cases, and test scripts.
Create and maintain production testing infrastructure and availability reporting.
Collaborate with cross-functional engineering teams to understand system functions and interactions.
Drive improvements in software development testing processes.
Provide clear and concise feedback to engineering and product teams.
Measure uptime metrics like correctness, availability, and latency SLOs/SLIs.
Requirements
Proven track record as a software engineer or similar role.
Programming experience with Go, Rust, or Python.
Experience designing, implementing, and maintaining secure and highly-available distributed systems.
Ability to develop, document, and execute test and SLO plans, test cases, and test scripts.
Excellent communication skills.
Nice-to-Haves
Experience working with synthetic traffic and load testing tools.
Experience working with Clickhouse, Prometheus, GraphQL, Postgres.
Experience working with data pipelines with a focus on reliability and scale.
Software Engineer building and scaling the full-stack Insights practice management platform (React, Python, PostgreSQL) for RCM and healthcare operations at Commure. Focus on performance, reliability, observability and delivering high-impact features for thousands of practices.
Build and operate systems for provisioning, monitoring, and orchestrating Cloudflare's private network interconnects at massive scale. Requires systems programming (Rust/Go), deep networking knowledge (BGP, Layer 3), distributed systems experience, and on-call participation.
Salary not listed
Hybrid5+ YOEFullstack Engineering
Full-Stack Engineer – AI Agent Platform (Fraud, Risk & AML)
OscilarPalo Alto, CA
Build full-stack systems for AI agents handling fraud detection, risk scoring, and AML using Python, TypeScript, React, AWS, and LLM frameworks. Requires 4+ years experience in fast-paced environments with distributed systems expertise.
Salary not listed
Hybrid4+ YOEFullstack Engineering
Software Engineer, Networking
TailscaleUnited States
Software Engineer building and scaling Tailscale’s global networking infrastructure (Funnel, DERP relays, dataplane). Design networking features, troubleshoot complex connectivity issues, maintain global services with SRE/DevOps practices. Requires deep networking expertise, Go experience preferred, and comfort with distributed systems.
163k – 204k/yr
Remote5+ YOEFullstack Engineering
Software Engineer, Observability
VercelNew York, NY
Software Engineer on Vercel's Observability team building large-scale data ingestion, processing, and visualization tools that integrate with frontend frameworks to help users monitor application health and performance. Requires 5+ years experience with JS/TS/Go, columnar databases, and open-source contributions.