Software Engineer - Infrastructure
Builds and maintains infrastructure components for ML inference platform using Python and Go. Implements Kubernetes deployments, monitoring systems, and resource management for efficient model serving, requiring Kubernetes knowledge and ML basics.
About the job
Responsibilities
- Develop infrastructure components for our ML inference platform using Python and Go
- Implement and maintain Kubernetes deployments for model serving
- Contribute to our inference orchestration layer for model deployments
- Build and enhance monitoring systems for model performance metrics
- Implement efficient resource management solutions for ML workloads
- Support infrastructure automation to improve ML deployment workflows
- Work closely with team members to implement technical solutions
- Help balance performance optimization with system reliability
- Participate in technical discussions around infrastructure improvements
- Learn and apply infrastructure best practices
Requirements
- Bachelor's degree or higher in Computer Science or related field
- Proficient coding abilities in one or more popular programming or scripting languages; Go proficiency is a plus
- Working knowledge of Kubernetes and containerization
- Basic understanding of machine learning concepts and model serving
- Familiarity with distributed systems concepts
- Experience with basic monitoring and logging tools
- Interest in ML/AI infrastructure and willingness to learn
- Strong collaboration and communication skills
Skills
Python, Go, Kubernetes, Docker, Distributed Systems, Monitoring Tools, Ml Inference, Resource Management, Containerization, Prometheus
Similar jobs
DevOps / SRE jobsInfrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.