Sr. Member of Technical Staff
Develops resilient, high-availability software for AI inference on AWS, including deployment workflows, container orchestration with Docker/Kubernetes, monitoring, and debugging. Requires Master's in CS and 18 months experience with AWS services, IaC tools, and Python.
About the job
Job Duties
- Design and develop software features that support system resiliency and high availability, including automated recovery mechanisms and fault-tolerant architecture across distributed environments.
- Develop and maintain cloud-based deployment workflows for AI inference software using AWS tools and services to support low-latency and scalable system performance.
- Develop Python-based scripts and APIs to streamline data preprocessing, inference execution, and post-processing for real-time inference tasks.
- Use parallel programming techniques (e.g., multi-threading, asynchronous processing) to maximize resource efficiency on AWS compute instances.
- Develop software components to support visualization and analysis of system performance metrics, enhancing the monitoring and usability of inference services.
- Develop inference software in Docker containers and define Kubernetes orchestration strategies that ensure software reliability and efficient scaling.
- Develop automated scripts to detect and mitigate common failure modes, improving software system reliability.
- Debug issues related to model deployment, container orchestration, networking configurations, documenting steps to reproduce and root-cause defects.
- Triage and resolve defects in the software service by analyzing logs, metrics, and distributed traces using tools like AWS CloudWatch, Grafana, or custom Python scripts.
- Work with product management and user experience teams to define requirements for inference service interfaces, including configuration, monitoring, and event logging.
- Author detailed technical documentation for infrastructure configurations, inference workflows, and APIs, ensuring clarity for internal teams and external customers.
- Document and track defects, enhancements, and release notes using tools like Jira and Git, ensuring version control and traceability.
Minimum Requirements
- Master’s degree or foreign equivalent in Computer Science or related field.
- 18 months experience as Information Security Analyst, Software Engineer, Sr. Member of Technical Staff, IT Senior Applications Engineer, or related.
- Infrastructure-as-Code and deployment automation: Terraform, AWS CloudFormation, AWS CDK, Ansible.
- Containerization and orchestration: Docker, Kubernetes, AWS EKS, AWS ECS, AWS Fargate, Helm.
- Compute and serverless services: AWS EC2, AWS Lambda, Auto Scaling Groups.
- Monitoring, logging, distributed tracing: AWS CloudWatch, AWS X-Ray, ELK (Elasticsearch, Logstash, Kibana), Prometheus, Grafana.
- Programming languages and frameworks: Python, Node.js, JavaScript, Flask.
- Data storage and caching: PostgreSQL, Redis, NFS.
- CI/CD and version control: Jenkins, Git.
Compensation
- Salary: $230,000 - $250,000 per year
Skills
Python, Docker, Kubernetes, AWS, Terraform, Aws Cloudformation, Aws Cdk, Ansible, Aws Eks, Aws Ecs, Aws Fargate, Helm, Aws Ec2, AWS Lambda, Aws Cloudwatch
Similar jobs
DevOps / SRE jobsOwn and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.
Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.