Lead Scientific Imaging Systems Engineer
Leads end-to-end infrastructure for a scientific imaging platform, covering Linux administration, GPU/HPC systems, storage, upgrades, and vendor coordination. The role supports AI-enabled imaging workflows and requires extensive production Linux, Image Artist, GPU, HPC, and enterprise storage experience.
About the job
Responsibilities
- Standardize server configurations and Signals Image Artist deployments across sites, establishing documented production and development baselines.
- Plan and execute vendor-supported Image Artist upgrades for server and client components.
- Coordinate with vendors to resolve upgrade blockers and document release risks and benefits.
- Install, configure, and troubleshoot NVIDIA/CUDA driver stacks and GPU compute nodes.
- Support HPC cluster integration and resolve infrastructure compatibility issues.
- Enable AI-based imaging analysis tools, including Phenologic AI, and benchmark GPU versus non-GPU performance.
- Lead large-scale data migrations between environments.
- Troubleshoot iSCSI/LUN issues and NetApp storage connectivity while minimizing downtime and data loss.
- Apply operating-system patches, endpoint protection updates, and vendor mitigations across the fleet.
- Own vendor support tickets and escalations and serve as the technical liaison for platform status and upgrade planning.
- Maintain runbooks, upgrade documentation, and equipment inventories.
- Deliver regular reports on platform health and progress.
Requirements
- Extensive production or enterprise Linux administration experience, including SLES and/or RHEL.
- Direct experience installing, upgrading, and supporting Signals Image Artist or comparable high-content imaging analysis platforms.
- Experience managing client/server version compatibility.
- Working knowledge of NVIDIA/CUDA driver stacks, GPU hardware troubleshooting, and HPC scheduling concepts.
- Experience with enterprise storage protocols and diagnosing production server storage and network connectivity issues.
- Ability to manage and resolve technical issues through formal hardware and software vendor support channels.
- Documentation-driven approach to change management, including staged upgrade paths, risk assessments, and rollback planning.
Compensation and Benefits
- Base salary: $120,000–$145,000 USD annually.
- 100% employee medical, dental, and vision coverage.
- Paid time off and holidays.
- 401(k) match up to 5%.
- Educational benefits for career growth.
- Employee referral bonus.
- Flexible spending accounts for healthcare, parking, dependent care, and transportation.
Skills
Linux, Sles, Rhel, Signals Image Artist, Nvidia Cuda, Gpu Computing, Hpc, Hpc Scheduling, Iscsi, Netapp, Storage Networking, Change Management
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for designing and operating reliable, scalable production infrastructure, leading incident response, and improving observability and resilience. Requires 5+ years of reliability-focused engineering experience and expertise across cloud, infrastructure as code, Kubernetes, monitoring, and application development.
Senior network engineer responsible for designing, operating, and securing MongoDB’s global network and VPN infrastructure. The role requires 6+ years of networking or systems engineering experience, strong enterprise networking expertise, automation skills, and the ability to lead complex infrastructure initiatives.
Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.
Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.