This project provides a way to run PyTorch operations and keep tensors on a remote NVIDIA GPU while the application runs on a local client machine (for example, a Mac without CUDA). It exposes two integration paths: a PyTorch device that transparently places tensors on a remote GPU for programs that opt in, and a CUDA shim that implements libcuda, the CUDA runtime and key libraries (cuBLAS, cuDNN, etc.) so existing Linux CUDA binaries can use a remote GPU without recompilation. The simpler PyTorch device is intended for new or adapted PyTorch workloads; the CUDA shim targets compatibility with existing CUDA programs. A command-line helper opens an SSH tunnel and configures the connection; a minimal example shows creating a tensor on the remote device and performing operations locally while execution happens remotely.
The repository includes client and server implementations, shared protocol metadata, a Python package providing the device and launcher, tests, build scripts for the C++ client and CUDA shim, and product documentation with quickstarts, examples (including training and a nanoGPT example), performance notes and deployment guidance. Neither protocol authenticates or encrypts by default, so the recommended deployment uses SSH tunnels and firewall rules (the CUDA server listens on port 9713 by default) to restrict access. Source is Apache-2.0 licensed and includes generated C++ artifacts and development tooling for testing and building.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.