Skip to content
CognitionCognition

Software Engineer, Infrastructure

Build and operate the compute, orchestration, networking, developer platform, and reliability systems underlying AI agents and developer tools. The role requires large-scale infrastructure experience, Kubernetes and cloud expertise, Python proficiency, and a strong security and observability mindset.

About the job

Responsibilities

  • Design and operate sandboxed compute environments powering agent task execution, including VM orchestration, container management, resource scheduling, and isolation at scale.
  • Build internal infrastructure for CI/CD pipelines, deployment systems, developer tooling, and platform abstractions.
  • Define SLOs and build monitoring and alerting systems.
  • Lead incident response and postmortems to improve reliability.
  • Anticipate capacity and architecture needs as the product scales.
  • Partner with product engineers and researchers to design infrastructure for their systems.

Requirements

  • Experience designing and operating large-scale distributed infrastructure.
  • Hands-on proficiency with Kubernetes, cloud platforms such as AWS, Google Cloud, or Azure, and infrastructure-as-code tools such as Terraform.
  • Strong software engineering fundamentals and proficiency in Python.
  • Experience instrumenting systems, building dashboards, and designing effective alerts.
  • Experience with sandboxing, network isolation, and secure multi-tenant compute.
  • Prior experience at a frontier AI lab, applied AI company, or developer tools company.
  • BS, MS, or equivalent in Computer Science, Mathematics, Engineering, or a related technical discipline from a highly selective program.

Compensation & Benefits

  • Base salary: $260,000–$300,000, plus significant early-stage equity.
  • Fully paid medical, dental, and vision coverage for employees and dependents.
  • 401(k) with company match.
  • Private chef, snacks, and other perks.

Skills

Kubernetes, AWS, GCP, Microsoft Azure, Terraform, Python, Distributed Systems, Vm Orchestration, Containers, CI/CD, Observability, Sandboxing, Network Isolation, Incident Response, SLOs

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff - Reliability Engineering
$240k+/yrHybrid5+ YOEDevOps / SRE

Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.

OpenAI

OpenAI

San Francisco, CA

Systems Integration Engineer, Build Systems | Consumer Devices
$293k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.

OpenAI

OpenAI

San Francisco, CA

Network Engineer
$293k+/yrHybridDevOps / SRE

Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.