Infrastructure Engineer
Builds and maintains scalable AWS infrastructure, CI/CD pipelines with GitHub, and secure AI Ops platform for clinical AI/ML products. Requires 5+ years in DevOps/SRE, Kubernetes proficiency, Terraform expertise, and healthcare compliance experience.
About the job
Responsibilities
- Design cost-optimized, fault-tolerant infrastructure for scale: Propose enhancements to infrastructure design to expand client base, deploy new products, manage cloud costs, and ensure reliability.
- Streamline development and deployment: Define branching and promotion strategy for regulatory compliance. Build and maintain CI/CD pipelines using GitHub for automated testing and deployment.
- Establish and evangelize infrastructure best practices: Create guidelines and templates such as Terraform modules, and educate team members.
- Infrastructure support and maintenance: Continuous monitoring of system performance and reliability, apply software upgrades. Collaborate on troubleshooting and performance optimization.
- Secure infrastructure: Partner with SecOps to implement security best practices complying with HIPAA, HITRUST, FDA, and client requirements.
- AI Ops Platform Architecture: Architect and build a secure, internal AI Ops platform to host and manage AI/ML agents for infrastructure and DevOps optimization.
Minimum Qualifications
- 5+ years of experience building and operating production cloud infrastructure on AWS as a DevOps, Infrastructure, Site Reliability Engineer, or similar role.
- Proficient with Kubernetes, preferably with EKS, including cluster bootstrapping and day-2 ops.
- Strong operational knowledge of relational databases such as PostgreSQL/MySQL (backups, failover, performance tuning).
- Deep expertise in Terraform (or equivalent IaC) and building clean, scalable modules.
- Familiarity with observability tools, particularly Datadog.
- Experience building infrastructure with sensitive data containing PHI/PII.
- Knowledge of CI/CD pipelines, preferably with CircleCI.
- Excellent communication skills and ability to collaborate with cross-functional teams.
- Experience handling ambiguity in a startup.
Preferred Qualifications
- Experience using AI agents to optimize infrastructure management or DevOps workflows.
- Experience with disaster recovery or business continuity plans.
- Experience with multi-account, multi-cluster topologies.
- Experience building systems in healthcare, life sciences, or regulated industries.
- Chaos engineering or game-day facilitation.
- Experience implementing and maintaining a GitOps framework.
Skills
AWS, Kubernetes, EKS, Terraform, Postgres, MySQL, Datadog, CI/CD, CircleCI, GitHub, GitOps, HIPAA, Hitrust
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.