Pushing the bf16 multiplication clock frequency to the max on the nominal corner on IHP 130nm 5L node.
-
Updated
Mar 29, 2026 - Python
Pushing the bf16 multiplication clock frequency to the max on the nominal corner on IHP 130nm 5L node.
A LLaMA2-7b chatbot with memory running on CPU, and optimized using smooth quantization, 4-bit quantization or Intel® Extension For PyTorch with bfloat16.
A JAX implementation of stochastic addition.
A Pytorch implementation of stochastic addition.
Drop-in exact bf16 flash-attention for CUDA with a deterministic backward, tuned for Blackwell (sm_120 / RTX 5090).
In bfloat16 a trainer gives the same token a different log probability depending on its batch shape; measured on 8 models, with the controls that survived
AWS deployment stack for Gemma 3 on SageMaker with HuggingFace TGI, OpenAI-compatible API (Lambda + API Gateway), and OpenWebUI chat interface
English Whisper Turbo transcription on Google Colab TPU in 3 cells. JAX/XLA, reusable model worker, TXT/SRT/VTT, resumable audio, and clear error logs. No API key.
FP64-accurate linear algebra on BF16/FP32 hardware (TPU) via the Ozaki scheme, in JAX.
Empirically calibrated false-positive bound for one-sided checksum ABFT on bfloat16 GEMM, where the classical τ = Ku threshold is undefined. Catches 76% of exponent-bit flips against a pre-registered 90% bar.
Generate Spike extensions, assembly tests, SVA assertions & docs for custom RISC-V AI vector instructions from YAML specs. Bit-accurate FP8/BF16/INT4 numerics.
BF16 x BF16 -> FP32 multiplier and two-stage FP32 adder in synthesizable SystemVerilog, bit-exact against two independent reference models
Float accumulation order alone flips RL reward verdicts and sampler/trainer probabilities — reproduce it on real GPT-2, then remove it with an order-independent reduction. numpy-only, runs in seconds.
Modly extension for research-only Cube3D INT4 text-to-mesh generation with UI-managed weights and GLB output.
Experimental research on bit flips in language-model weights: from faults nobody chose to faults an attacker did.
Fast FP32-to-BF16 stochastic rounding for PyTorch and Triton, with BF16 AdamW and momentum SGD optimizer states.
A research toolbox for running and studying LLMs on constrained hardware, with memory, placement, inference, and benchmarking methods.
To associate your repository with the bfloat16 topic, visit your repo's landing page and select "manage topics."