Startups

Beam

GPU Cluster Infrastructure Engineer

New York, NY, US / San Francisco, CA, US / Remote (US)

Contract

**Beam** is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

# **About the Role**

We're building out our own GPU capacity and we're looking for an experienced contractor to help us stand up high-performance GPU clusters. The work runs from design review through bring-in, and you'll leave behind the operational foundation our team needs to run them.

  • Review cluster designs and bills of materials across compute, networking, and storage, and catch gaps before hardware is ordered.
  • Lead acceptance testing: validate cabling and optics, bring up the InfiniBand fabric, run burn-in, and hold vendors to their deliverables.
  • Stand up and validate high-performance storage alongside vendor teams.
  • Build the out-of-band management layer and firmware baselines, and secure the management plane for customer-facing environments.
  • Integrate hardware, fabric, and storage telemetry into our observability stack, with alerting and automated health checks.
  • Write runbooks, as-builts, and remote-hands procedures.
  • Provide escalation support after go-live and help our team ramp up.

**Skills & Experience**

  • You've built and operated NVIDIA HGX or DGX clusters in production at a GPU cloud, HPC center, or AI lab.
  • Hands-on experience with InfiniBand: subnet management and UFM, fabric bring-up, and diagnosing degraded links and optics. NDR or newer.
  • GPU node bring-up and burn-in: firmware, BMC/Redfish, DCGM, NCCL testing, PXE and imaging, and XID error triage.
  • Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.
  • Equally effective on the data center floor and remotely, including directing colo remote hands.
  • You troubleshoot methodically across hardware, fabric, and software, document as you go, and communicate clearly with technical and non-technical people.
  • Bonus: recent-generation NVIDIA platforms, bare-metal cloud operations, Ansible or similar automation, Prometheus/Grafana, NVIDIA certifications.

# Benefits

  • Competitive salary and meaningful equity
  • Join a fast-growing pre-series A company at the ground floor
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more

Sourced 2026-10-08