Training neural networks in TensorFlow 2.0 with 5x less memory
-
Updated
Feb 21, 2022 - Python
Training neural networks in TensorFlow 2.0 with 5x less memory
A Toolkit for Training, Tracking, Saving Models and Syncing Results
A memory profiler for NVIDIA GPUs to explore memory inefficiencies in GPU-accelerated applications.
A prefix-cache advisor for LLM serving infrastructure that recommends KV-cache capacity and eviction policies from your request traces/logs.
Demonstration of generating mini-batches in Tensorlfow from GPU memory.
Dynamic GPU Layer Swapping: Train large models on consumer GPUs with intelligent memory management
Phase-manifold registration of 4D-CT: full-resolution (1 mm) deformable registration of 144 million voxels in ~2.4-3.3 GB of GPU memory, with an executable verification suite.
Detailed VRAM profiler for transformer inference with per-layer breakdown, activation analysis, and a predictive memory model that predicts VRAM with <1.2% error. Shows that FFN layers dominate static memory and that measured runtime VRAM exceeds KV-cache estimates by 2-4x.
Event-driven benchmark of adaptive batch composition policies for LLM serving, measuring how prefill and decode interference affects TTFT, TPOT, and throughput under different memory pressure regimes.
Gradient checkpointing and VRAM auto-tuning for spiking neural network training. O(sqrt(T)) activation memory with bit-identical gradients.
Deadline-aware KV-cache scheduling for protecting decode-critical request-state under long-context LLM inference pressure.
CLI toolkit for LLM inference preflight, vLLM serving configuration, benchmarking, telemetry, and capacity analysis.
Research harness for evaluating query-time bounded elimination of reconstructable KV-cache witnesses in long-context transformer inference workloads. Related provisional filing: IN 202641062451.
Preflight checks and collision prevention for local AI inference workloads
A CLI tool for estimating GPU VRAM requirements for Hugging Face models, supporting various data types, parallelization strategies, and fine-tuning scenarios like LoRA.
Tiered GPU memory architecture for consumer AI inference. VRAM as execution cache, system RAM as passive staging layer.
Calibrated simulation benchmark for real-time LLM request routing, comparing complexity signals, output-length awareness, cost savings, and quality-risk trade-offs.
📊 A command line monitoring tool (graph) for NVIDIA GPUs
Research prototype for short-horizon working-set residency in Mixture-of-Experts inference
Event-driven benchmark of per-tenant KV-cache isolation policies for multi-tenant LLM serving, measuring SLO compliance, fairness, and stability under noisy neighbor and burst scenarios.
To associate your repository with the gpu-memory topic, visit your repo's landing page and select "manage topics."