Software Engineer, Infrastructure
Build and own the infrastructure platform supporting Topsort’s real-time auction engine, APIs, and developer tooling. The role requires production Kubernetes, cloud, infrastructure-as-code, CI/CD, distributed-systems, observability, and security experience.
About the job
Responsibilities
- Design, build, and maintain foundational infrastructure supporting compute, networking, deployment, and observability.
- Define and enforce SLOs, build alerting systems, and lead preventative postmortems.
- Build reliable CI/CD pipelines and deployment systems while automating release toil.
- Manage AWS/GCP cloud infrastructure with Terraform or equivalent infrastructure-as-code tools.
- Build internal tooling and platform capabilities that improve developer speed, safety, and autonomy.
- Implement infrastructure-level security practices, including secrets management, network policies, RBAC, and supply-chain integrity.
- Capacity-plan, load-test, and tune systems proactively.
- Participate in the on-call rotation and respond to, resolve, and document production incidents.
Requirements
- Bachelor's degree in Computer Science, engineering, or a related field.
- 2–5 years of software engineering experience, with meaningful infrastructure, platform, or SRE experience.
- Hands-on production experience with Kubernetes.
- Experience with a major cloud provider, preferably AWS, and infrastructure-as-code using Terraform or a similar tool.
- Strong distributed-systems fundamentals, including the CAP theorem, eventual consistency, and failure modes.
- Experience building or maintaining CI/CD pipelines with GitHub Actions, ArgoCD, or similar tools.
- Familiarity with observability tooling for metrics, logs, and traces.
- Security-conscious approach to blast radius, least privilege, and secrets rotation.
Nice-to-haves
- Experience with high-throughput, low-latency systems such as ad tech, fintech, or real-time bidding.
- Familiarity with service mesh, eBPF, or advanced Kubernetes networking.
- Experience with database infrastructure and data streaming.
- Track record of improving developer experience and reducing time to deploy.
- Experience at a high-growth startup where infrastructure scaled rapidly.
Skills
Kubernetes, AWS, GCP, Terraform, CI/CD, GitHub Actions, Argo CD, Distributed Systems, Observability, RBAC, Secrets Management, Service Mesh, Ebpf, Data Streaming, Database Infrastructure
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.
Operates and scales shared telemetry infrastructure spanning metrics, logs, traces, alerting, dashboards, and profiling. The role requires at least three years of production engineering experience, distributed-systems troubleshooting, Infrastructure as Code, container orchestration, incident response, and on-call participation.