Software Engineer, Cloud Infrastructure
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.
About the job
Responsibilities
- Design, provision, and maintain AWS infrastructure using infrastructure-as-code tools such as AWS CDK or Terraform.
- Build CI/CD and testing for applications, infrastructure, and ML pipelines with GitHub Actions, CodeBuild, and CodePipeline.
- Operate secure networking with VPCs, PrivateLink, and VPC endpoints; manage IAM, KMS, Secrets Manager, and audit logging.
- Deploy and operate model endpoints using AWS Bedrock and/or SageMaker, selecting appropriate compute services for inference workloads.
- Build application services that call LLMs through clean APIs, including streaming, batching, and backoff strategies.
- Implement prompt and tool execution flows with LangChain or similar frameworks.
- Design chunking, embedding, ingestion, and retrieval pipelines for documents, time series, and multimedia data.
- Operate vector search systems and tune recall, latency, and cost.
- Build knowledge bases and data synchronization from S3, Aurora, DynamoDB, and external sources.
- Create evaluation harnesses and monitor quality, latency, regressions, model telemetry, token usage, and costs.
- Implement guardrails, rate limits, fallbacks, provider routing, PII detection, redaction, access controls, and audit trails.
- Build ingestion and processing pipelines with data integrity, lineage, and cataloging.
- Manage infrastructure supporting edge devices, secure messaging, identity, and over-the-air updates.
- Support application and product teams integrating retrieval services, embeddings, and LLM chains.
- Optimize retrieval, caching, inference, model selection, quantization, instance selection, and autoscaling.
- Own work from design through production, including on-call and continuous improvement.
Requirements
- Production experience shipping or operating LLM-powered applications.
- Strong AWS experience, including VPC, IAM, KMS, CloudWatch, S3, ECS/EKS, Lambda, Step Functions, Bedrock, and SageMaker.
- Experience with RAG design, prompt versioning, and chain orchestration using LangChain or similar tools.
- Experience building ingestion and transformation pipelines in Python.
- Familiarity with Glue, Athena, EventBridge, and SQS.
- Understanding of least privilege, secrets management, network isolation, and compliance for sensitive data.
- Experience using quantitative evaluations, A/B testing, and live metrics.
- Strong communication skills and ability to explain technical tradeoffs across product, security, and application engineering.
Nice-to-Haves
- 4+ years working with serverless or container platforms on AWS.
- Experience with vector databases, OpenSearch, or pgvector at scale.
- Experience with Bedrock Guardrails, Knowledge Bases, or custom policy engines.
- Familiarity with GPU workloads, Triton Inference Server, or TensorRT-LLM.
- Experience with big data tools for large-scale processing and search.
- Aviation data or other safety-critical domain experience.
- DevOps or DevSecOps experience automating CI/CD for ML and application services.
Compensation and Benefits
- Annual salary: $135,000–$260,000.
- 100% of employee medical premiums covered and 25% of dependent premiums covered.
- Three weeks of paid time off and 13+ paid company holidays.
- 401(k) offered.
Skills
AWS, Aws Cdk, Terraform, GitHub Actions, Kubernetes, Docker, LangChain, Amazon Bedrock, Amazon Sagemaker, Python, Opensearch, Pgvector, Amazon S3, Amazon Vpc, Aws Iam
Similar jobs
DevOps / SRE jobsBuild and operate deployment platforms, automation, and developer tooling that make software releases safer, more reliable, and self-service. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience with production systems and cloud or distributed infrastructure.
Build and operate highly available infrastructure for an enterprise AI platform, spanning cloud systems, Kubernetes, automation, observability, and reliability engineering. Requires 5+ years of production infrastructure experience, strong Python or Go skills, and daily use of AI-assisted workflows.
Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.
The DevOps Engineer will build and operate reliable infrastructure, deployment workflows, and observability for data pipelines and AI/ML systems. The role requires at least three years of DevOps, SRE, or infrastructure experience plus strong cloud, Terraform, containerization, and MLOps expertise.