Efficiency shorts (2026-10-09)
Short items from the 2026-10-09 window that touch inference cost, compression or GPU work but do not need their own page.
Kernels and serving
- KernelAgent (Meta, PyTorchCon talk). A multi-agent harness that feeds GPU hardware-performance counters into a closed loop for Triton kernel optimization. Across all 100 KernelBench L1 tasks: 2.02x over kernels from its earlier versions and 1.56x average over default torch.compile (post). Talk abstract, not a paper. See GPU kernels.
- MXFP4 checkpoints for AMD. Red Hat AI released more MXFP4 (4-bit microscaling floats, small blocks sharing one scale) checkpoints, e.g. Qwen3.8-27B, for vLLM on AMD GPUs (model). Follows its expert-only NVFP4 release (10-06).
- llama.cpp RPC across heterogeneous devices. Georgi Gerganov highlighted ggml's RPC backend for splitting inference across mixed devices (RT); 10-08 saw MiMo 2.6 Flash split across an RTX 6000 and an M5 laptop.
- Disaggregated inference (prefill and decode on separate hardware) got an endorsement from Cerebras CEO Andrew Feldman (post); YC's next Paper Club is on inference systems (RT).
- Upscale AI launched Token Fabric, scale-up networking for tightly coupled accelerators, congratulated by SambaNova (post). No specs captured.
Tokens and pricing
- GPT-6 uses the GPT-5 tokenizer. Simon Willison's ttok 1.0 switched its default; an independent test found all seven GPT-5.5 to GPT-6 models count 44,794 tokens on 31 fixtures (ttok 1.0). Contrast Haiku 5.5's new tokenizer, which counts the same text as 25-30% more tokens (10-08). Per-token prices are only comparable within a tokenizer.
- Byteification (Nature). Ai2, Cambridge and Edinburgh convert subword LLMs to byte-level models with under 1% of a pretraining budget, about 49B tokens (AI Weekly Espresso). Byte models read code and DNA more precisely; the trade is longer sequences.
Generation-side compression
- SemanTok (Stability AI). Coarse-to-fine video tokens whose early tokens are made more semantic; a model using it matches or beats one over 3x its size (blog).
- GRACE (arXiv 2610.10524, HF 74 upvotes): generation-aware latent compression for video diffusion. Abstract only.
- LittleBit (Samsung) resurfaced on X: a 13B model under 1 GB, i.e. below 1 bit per weight. Accuracy cost unverified (RT).
Evaluation cost
- Agentic RAG evaluation budgets (arXiv 2610.05034): at about 34M model tokens, spending on more questions lowers standard error 33% versus five reads per question and 12.6% versus three trajectories. Temperature zero cuts answer disagreement from 14.3% to 3.4%.
Related
KV cache · Quantization · GPU kernels · Efficiency shorts 10-08