Skip to content
#

kernel-optimization

Here are 40 public repositories matching this topic...

Compiler MVP that detects Transformer fusion patterns, generates optimized CUDA kernels with WMMA Tensor Cores, and executes them on real GPU hardware — 10.5 TFLOPs on RTX 2070, correctness validated against PyTorch.

  • Updated Jul 26, 2026
  • Python

Measurement harness for the sliding window attention premium in the vLLM TPU Ragged Paged Attention v3 kernel: per layer decode cost, block size control, throughput, and goodput for Gemma 4 31B on TPU v6e.

  • Updated Jul 29, 2026
  • Python

LLM agents that generate, verify, and evolve Triton GPU kernels. Includes a reward-hack-resistant benchmarking harness with strict correctness verification and fresh-input evaluation. Achieves up to 174.7× over PyTorch eager and outperforms FlexAttention (1.48×) and SDPA (1.17×) on selected workloads.

  • Updated Aug 12, 2026
  • Python

Add this topic to your repo

To associate your repository with the kernel-optimization topic, visit your repo's landing page and select "manage topics."

Learn more