Skip to content
ReductoReducto

Infrastructure Engineer

Founding Infrastructure Engineer to architect and scale resilient systems for AI/ML workloads, implement monitoring/observability, and automate infrastructure. Requires 5+ years production experience, Python, Kubernetes, and strong reliability focus.

About the job

Core Responsibilities

  • Designing, building, and maintaining highly available, scalable infrastructure to support intensive AI/ML workloads and real-time model deployments.
  • Implementing robust monitoring, alerting, and observability systems to ensure system health, performance, and uptime across cloud and on-prem environments.
  • Debugging, optimizing, and automating infrastructure for fast iteration and rapid deployment cycles, focusing on both reliability and developer velocity.
  • Proactively identifying, investigating, and resolving incidents to minimize downtime and maintain world-class service levels for enterprise customers.
  • Collaborating closely with engineers, ML specialists, and founders to shape product, infrastructure, and security strategies.

Requirements

  • 5+ years of hands-on experience in building or supporting production-grade infrastructure and reliability processes for high-throughput systems.
  • Comfortable with Python or similar languages, and exceptional at working across cloud platforms, container orchestration (e.g., Kubernetes), networking, and storage technologies.
  • Build your own tools on the fly to diagnose, experiment, and address reliability problems—whether it's an internal dashboard or an automated remediation workflow.
  • Quantitative, hands-on approach to system operations, automation, and continuous improvement.
  • Extremely high bar for quality and always aim for robust solutions rather than quick fixes.

Nice-to-Haves

  • Prior experience founding a company or building products/infrastructure in early-stage, high-growth environments.
  • Experience automating incident management processes with LLMs/AI.
  • Passion for open-source and contributions to reliability communities.
  • Experience building or optimizing monitoring, incident response, or high-performance computing systems for demanding AI/ML, fintech, or enterprise clients.
  • Keeping up with the latest trends in cloud, observability, and SRE best practices.

Benefits

  • Unlimited PTO
  • Free daily lunch at the office
  • Reimbursed transportation
  • Generous health insurance (medical, dental, vision)
  • Up to $150/mo health and wellness budget
  • Parental leave

Skills

Python, Kubernetes, Cloud Platforms, Networking, Storage Technologies, Monitoring, Observability, Automation, Incident Response, Sre Practices

Cloudflare

Cloudflare

Austin, TX
Systems Engineer - Database Platform
$150k+/yrHybridDevOps / SRE

Build and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.

Fluidstack

Fluidstack

New York, NY
Infrastructure Deployment Engineer
$150k+/yrOn-site5+ YOEDevOps / SRE

Leads on-site deployment of data center physical infrastructure, managing contractors, performing QA/QC on fiber optics and cabling, and ensuring compliance with standards. Requires 5+ years experience, SME-level fiber optic expertise, bachelor's degree, and 40% travel readiness.

Trexquant

Trexquant

New York, NY

Python Engineer - Trade Operations
$150k+/yrOn-site3+ YOEDevOps / SRE

The Python Engineer will improve and operate trading systems, support integrations with asset classes and prime brokers, and handle monitoring, incidents, and performance optimization. The role requires 3+ years of experience, strong Python and Linux skills, and familiarity with market data and order-entry systems.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

SimplePractice

SimplePractice

United States

DevOps Engineer, Data & AI Platform
$144k+/yrOn-site3+ YOEDevOps / SRE

The DevOps Engineer will build and operate reliable infrastructure, deployment workflows, and observability for data pipelines and AI/ML systems. The role requires at least three years of DevOps, SRE, or infrastructure experience plus strong cloud, Terraform, containerization, and MLOps expertise.