Code for https://arxiv.org/pdf/2608.16157
-
Updated
Sep 7, 2026 - Python
Code for https://arxiv.org/pdf/2608.16157
Fine-tune and serve large Mixture-of-Experts models on one GPU, with 4-bit experts in GPU memory, RAM or SSD.
4-bit GPU kernels that compute directly on packed Mixture-of-Experts weights, for training and inference.
Plan and run Mixture-of-Experts fine-tuning on one GPU: check the machine, fit the model, and keep a report of each run.
Trace which MoE experts actually fire and pin the hot ones to GPU - ran Qwen3-30B-A3B at 5.5-7.2 tok/s on a 4GB GTX 1050 Ti
gpu-offload: transparent CPU-to-GPU offloading. Runs the BLAS/LAPACK calls of unmodified programs (NumPy, R, Octave, C, Fortran) on an NVIDIA GPU when a per-machine calibration says it is faster, with CPU fallback. LD_PRELOAD + cuBLAS/cuSOLVER.
To associate your repository with the gpu-offloading topic, visit your repo's landing page and select "manage topics."