Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
-
Updated
Mar 19, 2026 - Python
Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.
Agent-queryable ROCm kernel optimization knowledge base for AMD Instinct MI300/gfx942 and MI350/MI355X/gfx950, packaged for Codex CLI and Claude Code with merged-PR provenance, real-silicon validation, and a maintainer-controlled pull-request evidence pipeline.
Open reproduction of NVIDIA's AVO paper (arXiv:2603.24517): evolutionary search where an autonomous coding agent IS the variation operator — Vary(P)=Agent(P,K,f). Runs on the Claude Code or Codex session you already have.
Evidence-driven CUDA, CUTLASS, Triton and GPU workload optimization for ChatGPT · 使用 ChatGPT 驱动 GPU workload 性能优化
Custom AWS Transform agent that migrates PyTorch/Triton kernels to AWS Trainium NKI (@nki.jit) and compiles, numerically verifies, and profiles every candidate on a real Trainium device before opening a PR.
Noeris — autonomous kernel fusion discovery + Triton autotuning for LLM kernels and Gemma layer deeper fusion (A100/H100 wins).
Compiler MVP that detects Transformer fusion patterns, generates optimized CUDA kernels with WMMA Tensor Cores, and executes them on real GPU hardware — 10.5 TFLOPs on RTX 2070, correctness validated against PyTorch.
Reward-hardened evaluation for LLM-generated GPU kernels, built on KernelBench.
Automatic Triton kernel generation and optimization for Intel GPU, powered by Claude Code.
Measurement harness for the sliding window attention premium in the vLLM TPU Ragged Paged Attention v3 kernel: per layer decode cost, block size control, throughput, and goodput for Gemma 4 31B on TPU v6e.
LLM agents that generate, verify, and evolve Triton GPU kernels. Includes a reward-hack-resistant benchmarking harness with strict correctness verification and fresh-input evaluation. Achieves up to 174.7× over PyTorch eager and outperforms FlexAttention (1.48×) and SDPA (1.17×) on selected workloads.
Hydra-2P, KDA, and CADR: fused PyTorch/Triton kernels for routing across model depth and sequence time.
Native AscendC Mamba2 selective scan / SSD forward-backward custom operator for Huawei Ascend 910B3 and 950PR, with CANN, torch_npu, A100 benchmarks and msprof profiling.
Paged attention decode kernel for LLM inference, implemented in both Triton and CUDA C++. Split-K (FlashDecoding-style) parallelism, Nsight-driven optimization, benchmarked against FlashInfer.
Profiling-driven micro-architectural case study & custom Triton GPU operators for batch-1 edge vision inference on NVIDIA Ada Lovelace.
Custom CUDA kernels for AWQ 4-bit LLM weight dequantization. A simple learning project.
Fast Triton kernels for multi-scale deformable attention (MSDA) — the core operator behind Deformable DETR, DINO, and Mask2Former — plus a from-scratch course and hands-on labs
An autonomous agent framework for evidence-driven GPU kernel optimization—explore, profile, tune, verify, and record every branch.
To associate your repository with the kernel-optimization topic, visit your repo's landing page and select "manage topics."