High Performance Computing
Position Objective
The HPC Systems Engineer will design, deploy, and optimise high-performance computing (HPC) clusters and GPU-enabled servers to support compute-intensive workloads in a cutting-edge environment. This role focuses on delivering robust, secure, and high-performing systems to empower researchers, data scientists, and engineers, ensuring seamless operation of HPC infrastructure while aligning with organisational goals and compliance standards.
Job Description and Responsibilities.
-
Design, deploy, and maintain HPC clusters and GPU-enabled servers for compute-intensive workloads, ensuring high availability and performance.
-
Administer and troubleshoot Red Hat Enterprise Linux systems in production environments to ensure stability and uptime.
-
Develop and maintain automation scripts (e.g., Ansible, Bash, Python) for system provisioning, configuration management, and patching.
-
Manage job scheduling systems (e.g., Slurm, PBS, LSF) and parallel file systems (e.g., Lustre, GPFS, BeeGFS) for efficient workload distribution.
-
Configure and optimise GPU workloads using NVIDIA CUDA or ROCm in HPC/AI environments.
-
Collaborate with researchers, data scientists, and engineering teams to fine-tune workloads and enhance system performance.
-
Ensure system security, stability, and compliance with organisational policies through proactive measures.
-
Monitor system health, analyse performance metrics, and conduct capacity planning to meet future demands.
-
Provide technical support, comprehensive documentation, and user training for HPC/GPU systems to enable seamless user experience
Qualifications & Skills
-
Bachelor’s degree in Computer Science, Information Technology, or a related field.
-
Minimum of 5-7 years of experience in HPC systems administration or related roles, with expertise in Red Hat Enterprise Lin
-
Proficiency in automation scripting (Ansible, Bash, Python) and managing job schedulers (Slurm, PBS, LSF) and parallel file systems (Lustre, GPFS, BeeGFs)
-
Hands-on experience with GPU workload configuration (NVIDIA CUDA, ROCm) in HPC/AI environments
-
Red Hat Certified Engineer (RHCE) or higher certification is preferred.
-
Experience in academic, research, or enterprise-scale HPC environments, with exposure to cloud-based HPC platforms (e.g., AWS ParallelCluster, Azure CycleCloud)
-
Familiarity with AI/ML workflows and tools (e.g., TensorFlow, PyTorch) in GPU environments is a plus.
-
Strong problem-solving, collaboration, and communication skills to support cross-functional teams.