Save application

alphagrepsecurities

Jun 09

HPC Engineer

Onsite Shanghai

python systems & network

greenhouse

Match & cover letter

Create a profile to see your match and get a cover letter for this role.

Create profile and match

Job description

You will design and optimize a large-scale GPU compute environment for high-performance workloads in Shanghai. This role involves ensuring the performance and reliability of the GPU cluster.

Details

  • Shanghai, work from office
  • Experience operating large GPU or HPC clusters in production
  • Skills required: Python, Bash, RDMA programming, Slurm, HPC networking, Linux system administration

The work

  • Design and operate a large-scale GPU compute environment
  • Identify and resolve performance bottlenecks
  • Manage workload schedulers like Slurm
  • Set up and manage high-performance shared storage systems