Save application

liquid-ai

Jul 28

Member of Technical Staff - GPU Infrastructure Engineer

Remote Remote · San Francisco

kubernetesgo research & engineeringfulltime

ashby

Match & cover letter

Create a profile to see your match and get a cover letter for this role.

Create profile and match

Job description

You will ensure the reliability of GPU clusters that support AI research and training. Your work will focus on improving resource efficiency and building automation tools.

Details

  • San Francisco, remote
  • Competitive salary with equity
  • Strong software engineering experience required
  • Experience with distributed systems, Linux, networking, and storage

The work

  • Own the reliability and operation of GPU clusters
  • Debug issues across compute, storage, networking, and distributed workloads
  • Improve resource utilization through tooling and automation
  • Onboard and migrate workloads across GPU providers and hardware platforms