A curated list for Efficient Large Language Models
-
Updated
Jun 17, 2025 - Python
A curated list for Efficient Large Language Models
[NeurIPS 2023 Spotlight] This project is the official implementation of our accepted NeurIPS 2023 (spotlight) paper QuantSR: Accurate Low-bit Quantization for Efficient Image Super-Resolution.
[ICML 2023] This project is the official implementation of our accepted ICML 2023 paper BiBench: Benchmarking and Analyzing Network Binarization.
The official implementation of the ICML 2023 paper OFQ-ViT
[ICLR 2026] This is the official PyTorch implementation of "QVGen: Pushing the Limit of Quantized Video Generative Models".
Chat to LLaMa 2 that also provides responses with reference documents over vector database. Locally available model using GPTQ 4bit quantization.
离线可用的本地类型化决策:4 核 CPU 单题 15.6ms。Local & offline Jev / System One inference on CPU — ONNX + INT8, no torch at runtime. 支持 laya / kev / PlayJev
A tutorial of model quantization using TensorFlow
Symmetric INT8 per-tensor and per-channel quantization engine for LLM weights and activations, with SNR and MSE reconstruction telemetry.
Symmetric INT8 per-tensor and per-channel quantization engine for LLM weights and activations, with SNR and MSE reconstruction telemetry.
Official ICML 2026 Spotlight implementation for structural MoE compression, including attribution-guided channel scoring, coverage-maximized pruning, compact checkpoint construction, and fine-tuning support.
PyTorch implementation of "BiDense: Binarization for Dense Prediction," A binary neural network for dense prediction tasks.
Measure quantization quality loss on Apple Silicon MLX — KL divergence, top-token flip rate and perplexity delta for KV-cache and weight quantization
This project distills a ViT model into a compact CNN, reducing its size to 1.24MB with minimal accuracy loss. ONNXRuntime with CUDA boosts inference speed, while FastAPI and Docker simplify deployment.
Linear Quantization From Scratch: PyTorch Implementation
Native mixed-precision quantization for ComfyUI models and text encoders W4A8, W4A4, INT8, FP8, streaming conversion, architecture-aware policies, and validation.
Comprehensive performance analysis of DeepSeek V3 quantization levels (FP16, Q8_0, Q4_0) on 16GB GPU environments.
PulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion Transformers
On-device Perceive → Reason pipeline for Apple Silicon: Core ML + Vision for perception, a swappable LanguageModel (Apple Foundation Models or Claude) for reasoning. Python conversion/quantization toolkit plus a SwiftUI reference app.
Pre-flight assurance for quantized conformal models — tells you what compression did to your coverage guarantee before deployment.
To associate your repository with the model-quantization topic, visit your repo's landing page and select "manage topics."