Running large language models on a single GPU for throughput-oriented scenarios.
-
Updated
Oct 28, 2024 - Python
Running large language models on a single GPU for throughput-oriented scenarios.
Run Mixtral-8x7B models in Colab or consumer desktops
A QoE-Oriented Computation Offloading Algorithm based on Deep Reinforcement Learning (DRL) for Mobile Edge Computing (MEC) | This algorithm captures the dynamics of the MEC environment by integrating the Dueling Double Deep Q-Network (D3QN) model with Long Short-Term Memory (LSTM) networks.
ZO2 (Zeroth-Order Offloading): Full Parameter Fine-Tuning 175B LLMs with 18GB GPU Memory [COLM2025]
LLM Inference on consumer devices
Monero hardware wallet protocol implementation for Trezor, agent
A framework for IoT devices to offload tasks to the cloud, resulting in efficient computation and decreased cloud costs.
A lightweight framework that enables serverless users to reduce their bills by harvesting non-serverless compute resources such as their VMs, on-premise servers, or personal computers.
30 tok/s for 20B MoE on 8 GB VRAM. Flat throughput to 32K context. Native MXFP4 + GGUF Q4_K/Q5_K/Q6_K via ggml CUDA kernels — zero dequant. Expert offloading for models that don't fit in GPU memory.
Code for paper "Real-time Neural Network Inference on Extremely Weak Devices: Agile Offloading with Explainable AI" (MobiCom'22)
Deep Reinforcement Learning for Energy-Efficient Computation Offloading in Mobile-Edge Computing
Backend.AI Client Library for Python
An Activation Offloading Framework to SSDs for Faster Large Language Model Training
A Pandas-inspired data analysis project with lazy semantics and query-offloading to SQLite
Unofficial FreeToken fork: on one RTX 3060 12 GB, a 35B MoE at 250k of context or gpt-oss-120b; Flash-Next 125B on two. Half the RAM, image input. Runs on Turing: RTX 2060, RTX 20 series, sm_75.
MOSAIC: Mobility-Oriented Scheduling and Intelligent Resource Allocation for IoT
An empirical study of benchmarking LLM inference with KV cache offloading using vLLM and LMCache on NVIDIA GB200 with high-bandwidth NVLink-C2C .
Routing-aware MoE expert offloading for vLLM — run DeepSeek-V4 and other huge MoEs on GPUs that can't hold all the experts (plugin, no fork).
Out-of-core LoRA fine-tuning and expert-routing measurement for Kimi K3 (2.78T MoE) on a 7.6 GB laptop: 93 layers streamed from a USB disk, forward checked against an independent C implementation, gradients checked by finite differences
To associate your repository with the offloading topic, visit your repo's landing page and select "manage topics."