Wiki Pages
Click a topic to expand. Concept pages first, then summaries newest-to-oldest.
concept page
Agent Evaluation & Benchmarks
concept page
Agent Harness Engineering (Loop, Harness, Graph)
concept page
Agent Memory
concept page
GUI Agents
concept page
Multi-Agent Systems
concept page
Self-Evolving Agents
concept page
Tool Use & Function Calling
2026-09-03 · Tier 2
The Harness Advantage in Autonomous Red Teaming (Ken Huang / Ridge Security)
2026-09-03 · Tier 2
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
2026-09-03 · Tier 2
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills (DisCo)
2026-09-02 · Tier 2
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
2026-09-02 · Tier 2
EM²Mem: Event-Centric Multimodal Memory for Large Language Models
2026-09-02 · Tier 2
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
2026-08-31 · Tier 2
ContextPilot: Teaching Agents Proactive Context Management via Fine-grained RL
2026-08-31 · Tier 2
J-Zero: Unified Challenger-Solver-Judge Co-Evolution from Zero Data
2026-08-30 · Tier 2
Your Agents Are Not Time Aware
2026-08-29 · Tier 2
Multi-Agent Design Patterns: Architectural Topologies, Failure Modes, and Production Hardening
2026-08-28 · Tier 2
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
2026-08-28 · Tier 2
CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
2026-08-28 · Tier 2
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
2026-08-28 · Tier 2
Training Agents to Evolve with Their Harness: Harness-Aware Training (TaoLive)
2026-08-28 · Tier 2
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
2026-08-26 · Tier 2
AutoSaddler: harness optimization as offline learning from mini-batches of failure traces
2026-08-26 · Tier 2
Nine practical rules for agents doing real work (Gradient Flow)
2026-08-26 · Tier 2
Recuris: recursive experiential-working memory evolution for long-horizon harnesses
2026-08-25 · Tier 2
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
2026-08-25 · Tier 2
Meta-Harness: Code-Space Harness Optimization with Raw Trace Access
2026-08-25 · Tier 2
Prime Agent: A Self-Improving RLM Harness
2026-08-25 · Tier 2
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection
2026-08-25 · Tier 2
Thinkingbox: One Success Isn't Reliability
2026-08-16 · Tier 2
Measuring Autonomous AI Research: 153 runs, 18 frontier models, one speedrun
2026-08-16 · Tier 2
SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
2026-08-16 · Tier 2
Specification-first convergence with an AI coding agent: 717k lines, 189 files, $2,430
2026-08-14 · Tier 2
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
2026-08-14 · Tier 2
DarwinX: Evolving Agent Harnesses Through Natural Selection
2026-08-14 · Tier 2
Ken Huang: Harness Engineering as a design-pattern language
2026-08-13 · Tier 2
Agent Safety Should Be a Runtime Contract
2026-08-13 · Tier 2
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
2026-08-13 · Tier 2
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
2026-08-13 · Tier 2
The Kurate Skills Cluster: Six Papers, One Unit of Abstraction
2026-08-13 · Tier 2
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
2026-08-12 · Tier 2
Three benchmarks, one shape of failure: DSAgentBench, SPIEval, VibeLifeBench
2026-08-12 · Tier 2
ALTK-Evolve: the agent playbook should be delivered selectively, not injected whole
2026-08-12 · Tier 2
Co-Evolution in Agentic Systems: a three-stage taxonomy for shedding human design
2026-08-12 · Tier 2
Mendel Gödel Machine: self-modification from the archive, not from the last failure
2026-08-12 · Tier 2
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents
2026-08-11 · Tier 2
Agent Memory Distillation (AMD): Hierarchical Teacher Memory for Small Agents
2026-08-11 · Tier 2
Harness Evolution: Ouroboros, Evo-Bench, and A²E
2026-08-11 · Tier 2
RoMeRL: Reduced-Order Utility States for Self-Evolving Agent Memory
2026-08-11 · Tier 2
SWE-Bench ProMax: Multilingual Large-Scale Refactoring, and an Audit of SWE-bench Verified
2026-08-10 · Tier 2
ReASearch: The Optimizer Is the Agent
2026-08-10 · Tier 2
StreamArena and StreamMind: hour-scale streaming video agents
2026-08-07 · Tier 2
Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
2026-08-06 · Tier 2
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
2026-08-06 · Tier 2
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
2026-08-06 · Tier 2
OneDayAgent: A Long-Horizon Harness for Autonomous Agents
2026-08-06 · Tier 2
Shadow Evaluations: AI Agents Cannot Yet Do Open-Ended AI Research
2026-08-06 · Tier 2
SKILL-KD: Contrastive Skill Distillation for LLM Agents
2026-08-05 · Tier 2
Do Agents Actually Learn From Experience? ContinualSkillBench and PAST-Bench
2026-08-05 · Tier 2
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
2026-08-04 · Tier 2
AAPT: Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path
2026-08-04 · Tier 2
ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step
2026-08-04 · Tier 2
SWE-Touch: coding agents lose 7.7 points when a human edits the code mid-task
2026-08-03 · Tier 2
ACM: Agentic Context Management for Long Horizon Tasks
2026-08-03 · Tier 2
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
2026-08-02 · Tier 2
Agent authority stops being one number: an effect ceiling for evolving agents and per-field certification for tool calls
2026-08-02 · Tier 2
Efficiency Matters in Autonomous Research
2026-08-02 · Tier 2
Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory
2026-08-01 · Tier 2
MemTX: Transactional Belief Commit for Stateful Agent Memory
2026-07-31 · Tier 2
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
2026-07-31 · Tier 2
Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
2026-07-31 · Tier 2
Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability
2026-07-31 · Tier 2
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
2026-07-31 · Tier 2
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
2026-07-31 · Tier 2
MemHarness: Memory Is Reconstructed, Not Replayed
2026-07-31 · Tier 2
Metis: Memory Foundation Model
2026-07-30 · Tier 2
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
2026-07-30 · Tier 2
Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems
2026-07-30 · Tier 2
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
2026-07-30 · Tier 2
Can AI Agents Conduct Open-Ended AI Research? Early Evidence from Two Case Studies
2026-07-30 · Tier 2
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
2026-07-29 · Tier 2
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
2026-07-27 · Tier 2
Agentic Context Management: Context Is a Budget, Not a Database
2026-07-27 · Tier 2
PRO-LONG: Keep Everything, Then Grep It
2026-07-27 · Tier 2
Skill Self-Play: Skills as the Middle Ground Between Verifiable and Open-Ended
2026-07-25 · Tier 2
AREX: Verification as the Control Signal for a Self-Improving Research Agent
2026-07-25 · Tier 2
Experience Distillation: Bake an Agent's Trial-and-Error Into Its Weights Without Touching the Environment Again
2026-07-25 · Tier 2
OpenForgeRL: Train Agents Inside the Harness They Actually Run In
2026-07-22 · Tier 2
Agent Tooling Day: AgentDebugX and DataFlow-Harness
2026-07-21 · Tier 2
Cursor: Agent Swarms and the New Model Economics
2026-07-21 · Tier 2
Environment-free Synthetic Data Generation for API-Calling Agents
2026-07-07 · Tier 2
Claude Fable 5, Part 3: Memory Engineering
2026-06-18 · Tier 2
CEO-Bench: Can Agents Play the Long Game?
2026-06-18 · Tier 2
OmniAgent: Native Active Perception as Reasoning for Omni-Modal Understanding
2026-06-18 · Tier 2
SkillOpt: Treating Agent Skills as Trainable Artifacts (Microsoft)
2026-06-18 · Tier 2
Xcientist: Externalizing Research Synthesis and Validation in AI Scientists
2026-06-17 · Tier 2
OPD-Evolver: distilling the *ability to evolve* into an agent, not just its memories
2026-06-17 · Tier 2
ProCUA-SFT Technical Report
2026-06-16 · Tier 2
FastContext: A Dedicated Exploration Subagent for Coding Agents
2026-06-15 · Tier 2
APPO: Agentic Procedural Policy Optimization
2026-06-15 · Tier 2
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
2026-06-15 · Tier 2
LLM Agents Can See Code Repositories
2026-06-15 · Tier 2
MRAgent: Memory is Reconstructed, Not Retrieved — Graph Memory for LLM Agents
2026-06-14 · Tier 2
EurekAgent: Agent Environment Engineering Is All You Need for Autonomous Scientific Discovery
2026-06-14 · Tier 2
EvoArena + EvoMem: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
2026-06-14 · Tier 2
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
2026-06-14 · Tier 2
Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents
2026-06-14 · Tier 2
HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness
2026-06-13 · Tier 2
ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs
2026-06-13 · Tier 2
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement (TRACE)
2026-06-13 · Tier 2
WebChallenger: A Reliable and Efficient Generalist Web Agent
2026-06-11 · Tier 2
Agentic Environment Engineering for LLMs: A Survey
2026-06-11 · Tier 2
Arbor: Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
2026-06-11 · Tier 2
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
2026-06-11 · Tier 2
DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch
2026-06-10 · Tier 2
Agent Self-Improvement: SearchSwarm (delegation) + RHO (harness optimization)
2026-06-09 · Tier 2
Honest Lying: Memory Confabulation in Reflexive Agents
2026-06-09 · Tier 2
LatentSkill: In-Context Textual Skills → In-Weight Latent Skills
2026-06-08 · Tier 2
Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedback
2026-06-08 · Tier 2
Disentangling Agent Self-Evolution: Does a Stronger Model Make a Better Self-Evolving Agent?
2026-06-08 · Tier 2
HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems
2026-06-08 · Tier 2
OpenSkill: Open-World Self-Evolution for LLM Agents
2026-06-08 · Tier 2
Self-Revising Discovery Systems: A Category-Theoretic Account of Agentic Scientific Discovery
2026-06-08 · Tier 2
SIA: Self-Improving AI with Harness & Weight Updates
2026-06-08 · Tier 2
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
2026-06-08 · Tier 2
When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents (ToolMaze)
2026-06-07 · Tier 2
ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
2026-06-07 · Tier 2
InKH: Interaction-Native Knowledge Harness for Financial LLM Agents
2026-06-06 · Tier 2
AdaPlanBench: Evaluating Adaptive Planning under World and User Constraints
2026-06-06 · Tier 2
DataCOPE: Unsupervised Skill Discovery for Agentic Data Analysis
2026-06-05 · Tier 2
Rethinking Continual Experience Internalization for Self-Evolving LLM Agents
2026-06-05 · Tier 2
EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management
2026-06-05 · Tier 2
MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery
2026-06-05 · Tier 2
MMPO: Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents
2026-06-05 · Tier 2
SePO: Self-Evolving Prompt Agent for System Prompt Optimization
2026-06-04 · Tier 2
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization (DRIFT / TELBench)
2026-06-04 · Tier 2
MemTrain: Self-Supervised Context Memory Training
2026-06-04 · Tier 2
StreamMA: Streaming Communication in Multi-Agent Reasoning
2026-06-02 · Tier 2
Multi-Agent Computer Use (MACU)
2026-06-01 · Tier 2
Compiling Agentic Workflows into Weights
2026-05-31 · Tier 2
CoHyDE: Iterative Co-Training of LLM Rewriter and Dense Encoder for Tool Retrieval
2026-05-31 · Tier 2
RePoT: Recoverable Program-of-Thought via Checkpoint Repair
2026-05-30 · Tier 2
PANDO: Single-Rollout Online Skill Distillation for Web Agents
2026-05-30 · Tier 2
Skill0.5: Joint Skill Internalization and Utilization for OOD Agentic RL
2026-05-29 · Tier 2
AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios
2026-05-28 · Tier 2
AI Research Agents Narrow Scientific Exploration
2026-05-28 · Tier 2
AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
2026-05-28 · Tier 2
AXPO: Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
2026-05-28 · Tier 2
BES: Bidirectional Evolutionary Search for Self-Improving Language Models
2026-05-28 · Tier 2
HRBench: Benchmarking Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs
2026-05-28 · Tier 2
LearnWeak: Automated Domain Specialization for Small Computer-Use Agents
2026-05-28 · Tier 2
LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?
2026-05-28 · Tier 2
OmniVerifier-M1: Multimodal Meta-Verifier with Symbolic Recalibration
2026-05-28 · Tier 2
PEAM: Parametric Embodied Agent Memory through Contrastive Internalization
2026-05-28 · Tier 2
ResearchMath-14K: Scaling Research-Level Math via Agents
2026-05-28 · Tier 2
ScientistOne: Chain-of-Evidence for Verifiable Autonomous Research
2026-05-28 · Tier 2
Verus-SpecGym: Agentic Environment for Specification Autoformalization
2026-05-28 · Tier 2
VibeSearchBench: Long-Horizon Proactive Search
2026-05-27 · Tier 2
How Do AI Agents Spend Your Money? Token Consumption in Agentic Coding
2026-05-27 · Tier 2
AKBE: Agentic Knowledge Boundary Enhancement
2026-05-27 · Tier 2
DarkForest: Controlled-Communication Multi-Agent Coordination
2026-05-27 · Tier 2
MUSE-Autoskill: Lifecycle-Managed Self-Evolving Skills
2026-05-27 · Tier 2
SAM: State-Adaptive Memory for Long-Horizon Reasoning Agents
2026-05-27 · Tier 2
Scaling the Harness: System Scaling as the Next Bottleneck in Agentic AI
2026-05-26 · Tier 2
Foundation Protocol + Static Authorization: the agent coordination + governance pair
2026-05-26 · Tier 2
MemForest: Hierarchical Temporal Indexing for Agent Memory
2026-05-26 · Tier 2
QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks
2026-05-26 · Tier 2
SEAL: Synergistic Co-Evolution of Agents and Learning Environments
2026-05-26 · Tier 2
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
2026-05-25 · Tier 2
From Raw Experience to Skill Consumption: a lifecycle study of model-generated agent skills
2026-05-25 · Tier 2
SkillOpt: a deep-learning-style optimizer for agent skills
2026-05-24 · Tier 2
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
2026-05-23 · Tier 2
Code as Agent Harness
2026-05-23 · Tier 2
Compound Engineering vs gstack vs Karpathy Autoresearch vs Superpowers vs RSI
2026-05-23 · Tier 2
SR²AM: Efficient Agentic Reasoning Through Self-Regulated Simulative Planning
2026-05-23 · Tier 2
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
2026-05-21 · Tier 2
SaaSBench: Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering
2026-05-21 · Tier 2
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
2026-05-20 · Tier 2
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
2026-05-20 · Tier 2
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
2026-05-20 · Tier 2
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
2026-05-19 · Tier 2
AI for Auto-Research: Roadmap & User Guide
2026-05-18 · Tier 2
Look Before You Leap: Autonomous Exploration for LLM Agents
2026-05-18 · Tier 2
MMSkills: Multimodal Skills for General Visual Agents
2026-05-18 · Tier 2
PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control
2026-05-18 · Tier 2
Solvita: Agentic Evolution for Competitive Programming
2026-05-17 · Tier 2
LIFE Survey: Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems
2026-05-16 · Tier 2
FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale
2026-05-16 · Tier 2
SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks
2026-05-15 · Tier 2
Agent Memory Cluster: STALE + Preping + EvolveMem + MemEye + MemLens + BOOKMARKS
2026-05-15 · Tier 2
EvoEnv: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
2026-05-15 · Tier 2
Orchard: Open-Source Agentic Modeling Framework — 67.5% SWE-bench Verified at 30B
2026-05-15 · Tier 2
SDAR: Self-Distilled Agentic Reinforcement Learning
2026-05-15 · Tier 2
WildClawBench: Native-Runtime Long-Horizon Agent Benchmark — Claude Opus 4.7 Tops Out at 62.2%
2026-05-14 · Tier 2
AgentLens: the Lucky Pass problem in SWE-agent evaluation
2026-05-14 · Tier 2
Context Training with Active Information Seeking
2026-05-14 · Tier 2
Revisiting DAgger in the era of LLM agents
2026-05-14 · Tier 2
MAP: a Map-then-Act paradigm for long-horizon interactive agents
2026-05-13 · Tier 2
Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks
2026-05-13 · Tier 2
LLM Agents Already Know When to Call Tools — Even Without Reasoning (Probe&Prefill)
2026-05-13 · Tier 2
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
2026-05-13 · Tier 2
Useful Memories Become Faulty When Continuously Updated by LLMs
2026-05-12 · Tier 2
X-OmniClaw: Unified Mobile Agent for Multimodal Understanding and Interaction
2026-05-11 · Tier 2
AutoTTS: Agentic Discovery for Test-Time Scaling
2026-05-10 · Tier 2
Jiayi Weng: Learning Beyond Gradients
2026-05-10 · Tier 3
OncoAgent: Dual-tier multi-agent framework for privacy-preserving oncology decision support
2026-05-09 · Tier 2
AI Co-Mathematician
2026-05-09 · Tier 2
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
2026-05-09 · Tier 2
Beyond Semantic Similarity: Direct Corpus Interaction (DCI)
2026-05-09 · Tier 2
Skill Curation Cluster: StraTA, Skill1, SkillOS
2026-05-07 · Tier 2
BRIGHT-Pro and RTriever-4B: Reasoning-Intensive Retrieval for Agentic Search
2026-05-07 · Tier 2
MedSkillAudit: Domain-Specific Audit Framework for Medical Research Agent Skills
2026-05-07 · Tier 2
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
2026-05-05 · Tier 2
AcademiClaw: When Students Set Challenges for AI Agents
2026-05-05 · Tier 2
Ctx2Skill: From Context to Skills — Self-Evolving Multi-Agent Skill Extraction
2026-05-05 · Tier 2
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
2026-05-05 · Tier 2
T^2PO: Token- and Turn-Level Policy Optimization for Stable Multi-Turn Agentic RL
2026-05-04 · Tier 1
Why Your Agentic AI Pentester Is Probably Just a Fancy Scanner — Ken Huang
2026-05-04 · Tier 3
LWD — Learning While Deploying: Fleet-Scale RL for Generalist Robot Policies
2026-05-02 · Tier 1
Ken Huang Ch 15 — Structured Output and Schema-Constrained Generation
2026-05-01 · Tier 2
Ara: Agent-Native Research Artifacts
2026-05-01 · Tier 2
Claw-Eval-Live: Live Agent Benchmark for Evolving Real-World Workflows
2026-05-01 · Tier 2
Eywa: Heterogeneous Scientific Foundation Model Collaboration
2026-05-01 · Tier 2
InteractWeb-Bench: Multimodal Agents under Non-Expert User Instructions
2026-05-01 · Tier 2
Intern-Atlas: A Methodological Evolution Graph
2026-05-01 · Tier 2
MCP Integration: Claude Code vs Hermes (Ken Huang Ch 13)
2026-05-01 · Tier 2
Synthetic Computers at Scale: Long-Horizon Productivity Simulation
2026-04-30 · Tier 2
ClawGym: A Scalable Framework for Building Effective Claw Agents
2026-04-25 · Tier 2
Claude Code Memory Systems: Chapter 8 Analysis
2026-04-23 · Tier 2
Claude Code vs. Hermes Agent: Permission System Architectures
2026-04-23 · Tier 2
Persistent Agent Infrastructure: Kimi K2.6, OpenAI Agent Studio, Anthropic Conway
2026-04-22 · Tier 2
AgentSPEX: An Agent Specification and Execution Language
2026-04-22 · Tier 2
SimpleTES: Evaluation-Driven Scaling for Scientific Discovery
2026-04-22 · Tier 2
HuggingFace ml-intern: Open-Source Agentic Post-Training Loop
2026-04-21 · Tier 2
Precise Debugging Benchmark: Models Regenerate, They Don't Debug
2026-04-21 · Tier 2
Reward-Free Self-Evolution: Agents That Learn Without Being Told What to Learn
2026-04-20
GTA-2: Benchmarking General Tool Agents from Atomic Use to Open-Ended Workflows
2026-04-20
PRL-Bench: LLMs on Frontier Physics Research
2026-04-20
Chapter 3: The Query/Agent Loop — Claude Code vs. Hermes Agent
2026-04-19 · Tier 2
Claude Code Architecture: A Deep Reading
2026-04-19 · Tier 2
UniDoc-RL: RL-Based Visual RAG with Hierarchical Actions
2026-04-18 · Tier 2
Corpus2Skill: Don't Retrieve, Navigate
2026-04-18 · Tier 2
DR3-Eval: Realistic Benchmark for Deep Research Agents
2026-04-17 · Tier 2
Dive into Claude Code: Architecture Analysis
2026-04-17 · Tier 2
SuperLocalMemory V3.3: Biologically-Inspired Agent Memory
2026-04-16 · Tier 2
DefenseClaw, MAESTRO, and the Security Boundary Agentic AI Has Been Missing
2026-04-16 · Tier 2
Do AI Coding Agents Log Like Humans? An Empirical Study
2026-04-16 · Tier 2
Exploration and Exploitation Errors Are Measurable for Language Model Agents
2026-04-16 · Tier 2
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
2026-04-16 · Tier 2
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration
2026-04-16 · Tier 2
UI-Copilot: Long-Horizon GUI Automation via Tool-Integrated Policy Optimization
2026-04-16 · Tier 2
Inside VAKRA: Reasoning, Tool Use, and Failure Modes of Agents
2026-08-30 · Tier 3
Salesforce moves Agentforce billing from seats to outcomes
2026-08-28 · Tier 3
Nvidia Agrees to Buy Hugging Face for $12.9 Billion
2026-08-26 · Tier 3
Ramp's Inspect: a non-AI company builds its own harness and writes 75% of its PRs with it
2026-08-16 · Tier 3
Optima: Artificial Analysis ships cost-and-time-per-task benchmarking
2026-08-14 · Tier 3
Token Price Is Not Task Cost: The AlphaSense Study
2026-08-13 · Tier 3
Grok 4.6: frontier intelligence sold on steps-per-task, not on the benchmark
2026-08-13 · Tier 3
Interconnects: "I wrote an AI textbook. How long until AI can do it better?"
2026-08-12 · Tier 3
Continual learning is arriving in pieces, and the argument for it is a maintenance bill
2026-08-11 · Tier 3
Gary Marcus: Open-Source Is Not the Same as Open-Weight
2026-08-11 · Tier 3
NVIDIA Turns AI Compute Into an Investable Asset Class: $500B With Six Capital Giants
2026-08-11 · Tier 3
The Router Market Repriced: Stripe–OpenRouter at $10B and the Scramble Behind It
2026-08-10 · Tier 3
Characterizing the Quality Profile of AI-Generated C++ in Production
2026-08-06 · Tier 3
ByteDance's Founder Rules Out Distillation
2026-08-02 · Tier 3
Three open letters in five days: 235 companies for open weights, Anthropic against, 1,324 lab employees asking to be slowed down
2026-07-29 · Tier 3
How Building Software Is Changing at Anthropic
2026-07-29 · Tier 3
The Actual Reason Why Google "Fell Out" of the AI Race
2026-07-29 · Tier 3
The Big AI Labs Are Suddenly Competing with Your Own Data
2026-07-29 · Tier 3
Sorry, Sam and Elon, We Have Not Reached the Singularity
2026-07-25 · Tier 3
The Open-Weights Letter: NVIDIA, Meta, Microsoft, and the Coalition Against Restriction
2026-07-08 · Tier 3
Anthropic IPO: 3Q26 Profit Over $1B
2026-06-05 · Tier 3
Anthropic: "When AI builds itself" — recursive self-improvement and the case for a pause button
2026-05-30 · Tier 3
Salesforce: 231-Day Migration Done in 13 Days, 79% More PRs, 5% Fewer Incidents
2026-05-29 · Tier 1
Anthropic Opus 4.8, Dynamic Workflows, and the $65B Series H at $965B Valuation
2026-05-28 · Tier 3
SemiAnalysis: Finding Miscompiles for Fun, Not Profit
2026-05-28 · Tier 3
Simon Willison: Anthropic and OpenAI Have Found Product-Market Fit
2026-05-26 · Tier 3
Google DeepMind's AlphaProof Nexus solves decades-old math problems for a few hundred dollars
2026-05-23 · Tier 3
Alibaba Qwen3.7-Max: 35 hours of autonomous code optimization on a custom chip
2026-05-23 · Tier 3
Anthropic Project Glasswing: Claude Mythos Preview finds 10,000+ critical vulnerabilities in one month
2026-05-23 · Tier 3
Kimi K2.5 powers Cursor Composer 2.5 at 1/10th the cost via Fireworks
2026-05-23 · Tier 3
Qwen3.7-Max: Alibaba's 35-hour autonomous coding agent
2026-05-21 · Tier 3
Anthropic on track for first profitable quarter, on the back of a $15B/yr SpaceX compute deal
2026-05-17 · Tier 2
Open Artifacts #21: The May 2026 Open-Model Wave and the CAISI / ECI Gap
2026-05-13 · Tier 2
Anthropic overtakes OpenAI in B2B adoption for the first time (Ramp data)
2026-05-13 · Tier 2
Recursive emerges from stealth with $650M for self-improving AI
2026-05-10 · Tier 3
Broadcom won't build OpenAI's custom chip without Microsoft buying 40 percent
2026-05-08 · Tier 1
Anthropic ↔ Colossus 1 Deal: Capacity Crunch + Brand Risk
2026-05-08 · Tier 1
GitHub Reliability Crisis: AI Load Breaks the Platform
2026-05-08 · Tier 2
Lambert: Notes from inside China's AI labs
2026-05-04 · Tier 3
Anthropic + OpenAI Both Build Services Companies Around Their AI
2026-05-03 · Tier 3
Microsoft VS Code Auto-Inserts "Co-Authored-by Copilot" Even With AI Off
2026-05-03 · Tier 1
Xiaomi MiMo-V2.5-Pro — Open-Weight Long-Horizon Coding at 40-60% Fewer Tokens
2026-05-02 · Tier 3
ChatGPT Tracks Free Users for Ads by Default
2026-05-01 · Tier 3
AISN #72 — CAIS AI-Wellbeing Research, Public-Sentiment Decline, OpenAI Releases
2026-05-01 · Tier 3
Anthropic Launches Claude Security — Defensive Cyber Productization
2026-05-01 · Tier 3
Chinese AI Startups Onshoring: Moonshot, StepFun Dissolving Offshore Structures
2026-05-01 · Tier 3
UK AISI: GPT-5.5 Matches Claude Mythos on Full Network Attack Simulation
2026-05-01 · Tier 2
Marcus: "The Greatest Capital Misallocation in History?"
2026-05-01 · Tier 3
Pentagon Signs Eight Tech Giants for AI-First Fighting Force; Anthropic Excluded
2026-05-01 · Tier 2
Pragmatic Engineer: AI Load Breaks GitHub; Anthropic's Trust Speedrun
2026-05-01 · Tier 1
SemiAnalysis: AI Value Capture — The Shift to Model Labs
2026-04-30 · Tier 3
Zig's Anti-LLM Contribution Policy and the "Contributor Poker" Argument
2026-04-29 · Tier 3
Claude for Creative Work: MCP Connectors for Blender, Adobe, Ableton, Autodesk
2026-04-29 · Tier 3
UCP Wins the Agentic Commerce Governance Layer
2026-04-22 · Tier 3
Amazon-Anthropic $33B Deal and AI Capital Concentration Week
2026-04-18 · Tier 3
OpenAI Executive Departures and Product Restructuring (April 2026)
2026-04-17 · Tier 3
Anthropic's Mythos Model: Government Access and the Trust Debate
2026-04-17 · Tier 3
Claude's Explosive Market Share Surge (April 2026)
concept page
LLM Routing
2026-09-02 · Tier 1
Safin-1: Safety from Within through Memory-Native State Evolution (MARCH routing)
2026-08-25 · Tier 1
Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation
2026-08-14 · Tier 1
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
2026-08-11 · Tier 1
Macaron-V1: Mixture-of-LoRA and Recursive Model-Harness Co-Design
2026-08-10 · Tier 1
SMRC-SD: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
2026-08-06 · Tier 1
Google Cloud ships cross-vendor LLM routing as managed infrastructure
2026-08-05 · Tier 1
VI-MoLE: Value-of-Information Routing for Mixtures of LoRA Experts
2026-08-04 · Tier 1
Kilo: The State of Open-Weight Models in AI Code Review Workflows
2026-07-31 · Tier 1
Beyond Geometric Complementarity: Coherent Overlap in Sparse MoE Routing
2026-07-31 · Tier 1
Kilo Code: Open Weights Are 79% of the Workload, and a 25x Price Gap Buys 5 Points
2026-07-31 · Tier 1
Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
2026-07-27 · Tier 1
Cursor's Agent Swarm: Separate the Planner From the Workers
2026-07-27 · Tier 1
Multi-Head Latent Control: Reading the Router Off the Hidden States
2026-07-25 · Tier 1
Separating the Task from the Model — the 550x Cost Cut Behind a Fixed Contract
2026-07-25 · Tier 1
Microsoft MAI Routing, and Stripe's Reported $10B Bid for OpenRouter
2026-07-24 · Tier 1
Sakana Fugu Ultra v1.1 — Model Router Claims to Beat Fable 5
2026-07-20 · Tier 1
When Is Routing Meaningful? Diversity and Robustness in Language Model Societies
2026-07-15 · Tier 1
Model Routing Is Simple. Until It Isn't. (IBM Research)
2026-06-18 · Tier 1
SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks
2026-06-16 · Tier 1
Plan With the Strong Model, Implement With the Cheap One (Kilo)
2026-06-15 · Tier 1
Orchestra-o1: Omnimodal Agent Orchestration
2026-06-12 · Tier 1
VIA-SD: Verification via Intra-Model Routing for Speculative Decoding
2026-06-10 · Tier 1
DPVR: Dual-Path Vision Token Routing (Late-Layer Fusion)
2026-06-09 · Tier 1
Chiaroscuro Attention (CHIAR-Former): per-token operator routing by spectral entropy
2026-06-07 · Tier 1
Kilo Code Audit: MiniMax M3 vs Claude Opus 4.8, and the Case for Model-Task Routing
2026-06-04 · Tier 1
Perplexity Hybrid Orchestrator: Deciding What Runs Locally vs in the Cloud
2026-05-29 · Tier 1
When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systems
2026-05-25 · Tier 1
DAR: Diffusion-Adaptive Routing for cross-layer information flow in DiTs
2026-05-23 · Tier 1
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles
2026-05-17 · Tier 1
MoE-muP: How to Scale Mixture-of-Experts (From muP to the Maximally Scale-Stable Parameterization)
2026-05-16 · Tier 1
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
2026-05-15 · Tier 1
Dynamic Latent Routing: Joint Latent-Code and Routing Policy for LM Post-Training
2026-05-15 · Tier 1
RouteProfile: Elucidating the Design Space of LLM Profiles for Routing
2026-05-11 · Tier 1
CaRE: Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts
2026-05-11 · Tier 1
Conductor: Learning to Orchestrate Agents in Natural Language
2026-05-08 · Tier 1
State of Routing in Model Serving (Netflix Tech Blog)
2026-05-02 · Tier 1
Step-level Optimization for Efficient Computer-use Agents
2026-05-01 · Tier 1
Ken Huang Ch 14: Model Routing and Provider Abstraction (Claude Code vs Hermes)
2026-04-17 · Tier 1
TRACER: Trace-Based Adaptive Cost-Efficient Routing for LLM Classification
concept page
Compute Economics (GPU pricing, utilization, and the durability of a chip)
concept page
GPU Kernels and Accelerator Optimization
concept page
Memory Hierarchy for AI (concept)
2026-09-02 · Tier 1
The Physics of LLM Inference: Memory Walls, Arithmetic Intensity, and Compute Ceilings
2026-09-02 · Tier 1
TrainSDC: Characterizing and Mitigating Silent Data Corruption in LLM Training
2026-08-31 · Tier 1
Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization
2026-08-31 · Tier 1
SemiAnalysis: Most Neoclouds Suck At Security
2026-08-30 · Tier 1
Apple's Mac mini and Mac Studio become the local-inference escape hatch
2026-08-28 · Tier 1
AI Is Designing the Silicon: Hot Chips 2026
2026-08-28 · Tier 1
Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization
2026-08-26 · Tier 1
OpenAI Jalapeño: a first-generation inference ASIC that beats Rubin on tokens per megawatt
2026-08-12 · Tier 1
Semiconductor week 32, 2026: a record quarter, and two bets against HBM
2026-08-11 · Tier 1
TileRT: Compiling the Whole Decode Graph into One Persistent Kernel
2026-08-03 · Tier 1
JAXBench: Benchmarking Autonomous TPU Kernel Optimization
2026-07-30 · Tier 1
LEGO Datacenters: The Binding Constraint on AI Capacity Is Electricians
2026-07-27 · Tier 1
CXMT Opens at $487B: China's Memory Champion Gets a Public Balance Sheet
2026-07-25 · Tier 1
NVIDIA's $500B SK Partnership: Buying the Memory Supply, Not Just Selling Chips
2026-07-25 · Tier 1
Can AMD Break the CUDA Moat? (SemiAnalysis, Advancing AI 2026)
2026-07-23 · Tier 1
SLAI T-Rex: Full-Parameter Post-training of DeepSeek-V4 on Ascend SuperPOD
2026-07-22 · Tier 1
NVIDIA Vera Rubin Ramps to Gigascale: 10x Tokens per Megawatt
2026-07-22 · Tier 1
Meta's Infrastructure Team Needs a Culture Reset (SemiAnalysis)
2026-06-07 · Tier 1
Memory Technology for Agentic AI Workloads (Ken Huang)
2026-06-01 · Tier 1
NVIDIA GTC Taipei 2026: RTX Spark, Vera CPU, Nemotron 3 Ultra
2026-05-29 · Tier 1
AWS Resilient Network Graphs: A Flat Data-Center Fabric With 33% Higher Throughput and 40% Lower Network Power
2026-05-28 · Tier 1
AI-Generated CUDA Kernels Silently Break Training and Inference
2026-05-28 · Tier 1
DoubleAI Speed-of-Light Blackwell CUDA Kernels (SOL-ExecBench)
2026-05-26 · Tier 1
SemiAnalysis: Inside the 800VDC Revolution. Part 1
2026-05-21 · Tier 1
SemiAnalysis EDA Market Primer (Part 2): how chip-design software became a $16B/yr AI tailwind
2026-05-19 · Tier 1
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
2026-05-14 · Tier 1
Inference as energy-to-token production: a position paper
2026-05-13 · Tier 1
SemiAnalysis: Cerebras — Faster Tokens Please
2026-05-09 · Tier 1
KernelBench-X: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
2026-05-04 · Tier 1
Cerebras Targets $40B Valuation in Second IPO Attempt
2026-04-27 · Tier 1
Semiconductor Week 17, 2026: AI Memory Supercycle and Agentic EDA
2026-04-21 · Tier 1
SemiAnalysis: GPU Cluster Economics and the Goodput Reckoning
concept page
Knowledge Distillation
concept page
KV Cache
concept page
Model Pruning and Sparsity
concept page
Parametric Context Internalization
concept page
Speculative Decoding
concept page
Test-Time Compute Allocation
2026-09-03 · Tier 1
CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
2026-09-03 · Tier 1
Language Models Can Control Their Own Attention (Declarative Attention)
2026-09-03 · Tier 1
Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation
2026-09-02 · Tier 1
A Universal Context-Reuse Layer for Cross-Model KV Sharing
2026-09-02 · Tier 1
Functional Degeneracy in Neural Networks: Measurement and Pruning
2026-08-31 · Tier 1
When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs
2026-08-30 · Tier 1
FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning
2026-08-30 · Tier 1
Mapping the Concept Landscape: Structural Perception of Global Distributions for Transparent Data Pruning
2026-08-30 · Tier 1
Reservoir of Importance: Learning Semi-Structured Sparsity with Differentiable Subset Sampling
2026-08-29 · Tier 1
CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
2026-08-29 · Tier 1
KV vs Prefix vs Prompt vs Semantic Caching: the four layers, and which one can lie to you
2026-08-28 · Tier 1
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher
2026-08-28 · Tier 1
TTPO: Test-Time Policy Optimization
2026-08-26 · Tier 1
DiffusionOPSD: on-policy self-distillation turns image rewards into intermediate targets
2026-08-26 · Tier 1
OPDVR: On-policy Distillation with Verifiable Reward
2026-08-26 · Tier 1
OraRL: annotations as oracle rollouts, and the advantage-inversion failure
2026-08-26 · Tier 1
Quantization-Aware Healing: a 4-bit model that beats its own full-precision checkpoint
2026-08-25 · Tier 1
Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress (R2-OPD)
2026-08-25 · Tier 1
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration
2026-08-16 · Tier 1
AutoPrune: An AI4AI Framework for Visual Token Pruning
2026-08-16 · Tier 1
The Slow Death of Scaling, and Data Inside the Loop (Sara Hooker, Adaption)
2026-08-16 · Tier 1
CaRL: Knowing When to Quit — Diagnosing and Training LLMs to Abort Futile Reasoning
2026-08-16 · Tier 1
Gambit: Thought-Level Beam Search for Reasoning
2026-08-16 · Tier 1
Hinting: Self-Distillation Without a Golden Answer (Applied Compute)
2026-08-16 · Tier 1
Maglev: Sliding Recurrent Memory
2026-08-16 · Tier 1
The Teacher-Student Alignment Cluster: four papers, one diagnosis (2026-08-16)
2026-08-14 · Tier 1
DeepSeek Harness v0.1 and the price of a cache hit
2026-08-14 · Tier 1
Full-bandwidth Transformer: Latent Feedback Between Decoding Steps
2026-08-14 · Tier 1
LycheeMemory V2: Long-Term Agent Memory via Semantic Segment-Level Consolidation
2026-08-14 · Tier 1
Massive Activations in Hybrid Linear Attention LLMs: Pre-Attention Spikes and Inter-Spike Plateaus
2026-08-13 · Tier 1
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
2026-08-13 · Tier 1
NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs
2026-08-12 · Tier 1
From Sweep to Seam: interleaved cross-block post-training quantization
2026-08-11 · Tier 1
Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation
2026-08-11 · Tier 1
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
2026-08-10 · Tier 1
Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation
2026-08-10 · Tier 1
WorldTrace: Addressable Memory for Video World Models
2026-08-07 · Tier 1
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
2026-08-06 · Tier 1
OPD-V: Modality Balance as Privileged Information
2026-08-06 · Tier 1
Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation via a Pixel Bridge
2026-08-06 · Tier 1
RSTG: Recovering Learning Signals from Negative RL Groups via Adaptive Teacher Guidance
2026-08-06 · Tier 1
SA-OPD: Spurious-Signal-Aware On-Policy Distillation
2026-08-06 · Tier 1
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
2026-08-05 · Tier 1
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
2026-08-05 · Tier 1
Mixture-of-Kittens: Cursor's Open-Source MoE Training Megakernel for NVL72
2026-08-05 · Tier 1
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
2026-08-05 · Tier 1
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning
2026-08-05 · Tier 1
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
2026-08-04 · Tier 1
CRPO: Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
2026-08-04 · Tier 1
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
2026-08-03 · Tier 1
DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference
2026-08-03 · Tier 1
Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory
2026-08-03 · Tier 1
ReOPD: Multi-Turn On-Policy Distillation with Prefix Replay
2026-08-03 · Tier 1
VQVLA: Motion-Aware Vector Quantization with Centroid Reuse for VLA Inference
2026-08-02 · Tier 1
KAP: Knowledge Access Planning, or how the prompt format throws away everything the retriever knew
2026-08-02 · Tier 1
MAPD: distilling a closed teacher through a JSON protocol instead of its words
2026-08-01 · Tier 1
MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention
2026-08-01 · Tier 1
Speculative Decoding Leaves the Lab: OpenAI's 13x Price Move and tinygrad's 245 tok/s
2026-07-31 · Tier 1
β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
2026-07-31 · Tier 1
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
2026-07-31 · Tier 1
Flux-OPD: On-Policy Distillation with Evolving Contexts
2026-07-31 · Tier 1
Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes
2026-07-31 · Tier 1
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
2026-07-31 · Tier 1
OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
2026-07-30 · Tier 1
CAST: Game Solvers as Turn-Level Teachers for LLM Agents
2026-07-30 · Tier 1
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
2026-07-30 · Tier 1
Local Coding Models in 2026: KV Cache Is the Binding Constraint, Not Parameter Count
2026-07-30 · Tier 1
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
2026-07-29 · Tier 1
Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization (BPM)
2026-07-29 · Tier 1
LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
2026-07-29 · Tier 1
OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs
2026-07-29 · Tier 1
Pass the Baton: Trajectory-Relayed On-Policy Distillation (Relay-OPD)
2026-07-29 · Tier 1
Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV
2026-07-28 · Tier 1
A Frozen 12B Beats Frontier Models on Verified Work
2026-07-28 · Tier 1
Error Certificates for KV-Cache Eviction via Randomized Design
2026-07-27 · Tier 1
MXSens: Sensitivity-Aware Mixed-Precision Quantization
2026-07-27 · Tier 1
VisCo: The Model Is Already a Good Compressor of Its Own Vision Tokens
2026-07-26 · Tier 1
QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling
2026-07-25 · Tier 1
Requential Coding — Compressing a Model by Recording Only Where Teacher and Student Disagree
2026-07-25 · Tier 1
VCSD: Making a Model Its Own Teacher by Deleting the Image
2026-07-24 · Tier 1
Dataset Distillation by Influence Matching (Inf-Match)
2026-07-24 · Tier 1
Predictive Divergence Masks for LLM RL
2026-07-24 · Tier 1
ReOPD: Multi-Turn On-Policy Distillation with Prefix Replay
2026-07-24 · Tier 1
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals
2026-07-23 · Tier 1
FVAttn: Adaptive Sparse Attention with Runtime Load Balancing
2026-07-23 · Tier 1
RIPO: Beyond Euclidean Clipping (Riemannian Isometric Policy Optimization)
2026-07-23 · Tier 1
SLPO: Scaling Latent Reasoning via a Surrogate Policy
2026-07-22 · Tier 1
H²SD: Hybrid Hindsight Self-Distillation
2026-07-22 · Tier 1
ISO: An RLVR-Native Optimization Stack
2026-07-22 · Tier 1
Stale but Stable: Staleness-Adaptive Trust Regions for Asynchronous RL (SAT)
2026-07-21 · Tier 1
Distilled RL: Teacher-Guided Reinforcement Learning for LLM Post-training
2026-07-21 · Tier 1
FlashRT: Agent Harness for Deploying Real-Time Multimodal Applications
2026-07-21 · Tier 1
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune
2026-07-21 · Tier 1
TOPL: Token-Level Off-Policy Labeling for Faithful Generation
2026-07-17 · Tier 1
Byte-Exact KV-Cache Grafting: Verified Knowledge as a Reusable Cache Artifact
2026-07-17 · Tier 1
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
2026-06-18 · Tier 1
Quality-Aware OPSD: gate the teacher per coordinate-token by whether it can still reach the right box
2026-06-17 · Tier 1
d-OPSD: on-policy self-distillation for diffusion LLMs, via "self-future" suffixes
2026-06-17 · Tier 1
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
2026-06-17 · Tier 1
Rethinking the Role of Efficient Attention in Hybrid Architectures
2026-06-17 · Tier 1
Variable-Width Transformers
2026-06-17 · Tier 1
ZPPO: keep the teacher in the prompt, not the policy gradient
2026-06-16 · Tier 1
Tangram: Non-Uniform KV Cache Compression as a Serving Substrate
2026-06-16 · Tier 1
TokenPilot: Cache-Efficient Context Management for LLM Agents
2026-06-16 · Tier 1
VibeThinker-3B: Frontier Verifiable Reasoning in a 3B Model
2026-06-15 · Tier 1
AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization
2026-06-15 · Tier 1
Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation
2026-06-15 · Tier 1
LoRA-alpha: The Hidden Power of the Scaling Factor in LoRA Optimization
2026-06-15 · Tier 1
PoLar: Skip a Layer or Loop It? Learning Program-of-Layers in LLMs
2026-06-13 · Tier 1
See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents
2026-06-12 · Tier 1
Flash-GMM: a memory-efficient fused kernel for soft clustering
2026-06-12 · Tier 1
MiniMax Sparse Attention (MSA): the paper behind MiniMax-M3
2026-06-12 · Tier 1
SG-OPD: Sign-Gated On-Policy Distillation
2026-06-11 · Tier 1
Bebop: Breaking Entropy Bounds — accelerating RL training via MTP with rejection sampling
2026-06-11 · Tier 1
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning (M²LA KV compression)
2026-06-10 · Tier 1
Attention Amnesia in Hybrid LLMs: QK-Restore
2026-06-10 · Tier 1
DLA: Dynamic Linear Attention
2026-06-10 · Tier 1
Latent Memory: One Token per Multimodal Evidence
2026-06-09 · Tier 1
FlashMemory-DeepSeek-V4: Lookahead Sparse Attention (LSA)
2026-06-09 · Tier 1
On the Geometry of On-Policy Distillation
2026-06-09 · Tier 1
Latent Context Language Models (LCLMs): End-to-End Context Compression at Scale
2026-06-09 · Tier 1
Trajectory-Refined Distillation (TRD)
2026-06-08 · Tier 2
Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings (EmbedFilter)
2026-06-08 · Tier 1
RL-Kernel: A Low-Level Kernel Library for RLHF Training of LLMs
2026-06-07 · Tier 1
Flash-WAM: Modality-Aware Step Distillation for World-Action Models
2026-06-07 · Tier 1
SEAOTTER: Sensor-Embedded Autoencoding with One-Time Transcode
2026-06-06 · Tier 1
AdaCodec: A Predictive Visual Code for Video MLLMs
2026-06-06 · Tier 1
Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution
2026-06-06 · Tier 1
Video2LoRA: Parametric Video Internalization for Vision-Language Models
2026-06-05 · Tier 1
The Shadow Price of Reasoning: CLEAR and Economic Budget Allocation for LLMs
2026-06-05 · Tier 1
OPRD: On-Policy Representation Distillation
2026-06-04 · Tier 1
Echo-Infinity: Learnable Evolving Memory for Real-Time Infinite Video Generation
2026-06-04 · Tier 1
FiRe-OPD: Filter, Then Reweight On-Policy Distillation
2026-06-04 · Tier 1
Unlocking Feature Learning in Gated Delta Networks at Scale (μP for linear attention)
2026-06-04 · Tier 1
MergePipe: Budgeting Expert Reads for Weight-Space Model Merging
2026-06-04 · Tier 1
SDPG: Self-Distilled Policy Gradient
2026-06-04 · Tier 2
STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations
2026-06-03 · Tier 1
MiniMax M3 and Step 3.7 Flash: open-weight efficiency at the frontier
2026-06-03 · Tier 1
Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling
2026-06-03 · Tier 1
TrOPD: Trust Region On-Policy Distillation
2026-06-03 · Tier 1
VaSE: Value-Aware Stochastic KV Cache Eviction for Reasoning Models
2026-06-02 · Tier 1
Draft-OPD: On-Policy Distillation for Speculative Draft Models
2026-06-02 · Tier 1
κ-SwiGLU: Confidence-Adaptive SwiGLU for Mixture-of-Experts
2026-06-02 · Tier 1
LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning
2026-06-02 · Tier 1
LVSA: Training-Free Sparse Attention for Long Video Diffusion
2026-06-02 · Tier 1
Speculative Pipeline Decoding (SPD): Zero-Bubble Speculation via Pipeline Parallelism
2026-06-02 · Tier 1
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
2026-06-01 · Tier 1
dMoE: dLLMs with Learnable Block Experts
2026-06-01 · Tier 1
StateKV: Linear Scaling Video VLMs for Long Video Understanding
2026-06-01 · Tier 1
TA-OPD: Token Teachability in On-Policy Distillation
2026-05-30 · Tier 1
Conf-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage
2026-05-30 · Tier 1
EarlyTom: Early Token Compression Inside the Vision Encoder
2026-05-29 · Tier 1
How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
2026-05-29 · Tier 1
Parallax: Parameterized Local Linear Attention for Language Modeling
2026-05-28 · Tier 1
Less is More: Early Stopping Rollout for On-Policy Distillation (ESR)
2026-05-28 · Tier 1
OSP-Next: Sparse Sequence Parallelism, HiF8, and RL for Efficient Video Generation
2026-05-27 · Tier 2
CPT: Collaborative Parallel Thinking for Efficient Test-Time Scaling
2026-05-27 · Tier 1
The Efficiency Frontier: Cost-Performance Optimization in LLM Context Management
2026-05-27 · Tier 1
Language Models Need Sleep
2026-05-27 · Tier 1
MobileMoE: Scaling On-Device Mixture of Experts
2026-05-27 · Tier 1
RT-Lynx: Activation Sparsity for Diffusion Transformers
2026-05-26 · Tier 1
Channel-wise Vector Quantization (CVQ)
2026-05-26 · Tier 1
SMART: Your Embedding Model is Smarter Than You Think
2026-05-25 · Tier 1
BitCPM-CANN: native 1.58-bit training outside the CUDA ecosystem
2026-05-25 · Tier 1
DeepSeek-V4: interleaved compressed attention with manifold-constrained hyper-connections
2026-05-25 · Tier 1
Good Token Hunting: training-free token selection for Visual Geometry Transformers
2026-05-25 · Tier 1
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
2026-05-25 · Tier 1
Pion: a high-pass replacement for Muon outside pretraining
2026-05-24 · Tier 1
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
2026-05-24 · Tier 1
KVServe: Service-Aware KV Cache Compression for Disaggregated LLM Serving
2026-05-24 · Tier 1
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps (RTPurbo)
2026-05-24 · Tier 1
WorldKV: Efficient World Memory with World Retrieval and Compression
2026-05-23 · Tier 1
Forecasting Downstream Performance of LLMs With Proxy Metrics
2026-05-22 · Tier 1
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
2026-05-22 · Tier 1
KVServe: Service-Aware KV Cache Compression for Disaggregated LLM Serving
2026-05-22 · Tier 1
RTPurbo / Full Attention Strikes Back: Sparse Attention Transfer in Hundreds of Steps
2026-05-22 · Tier 1
WorldKV: Training-Free World Memory via KV Chunk Retrieval and Compression
2026-05-21 · Tier 1
110 tok/s on 12GB VRAM with Qwen3.6-35B-A3B via ik_llama.cpp and MTP
2026-05-21 · Tier 1
Mix-Quant: Phase-Aware NVFP4 for Agentic LLM Prefill, BF16 for Decoding
2026-05-21 · Tier 1
How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
2026-05-21 · Tier 1
OCTOPUS: Joint Triplet Quantization for KV Cache via Octahedral Parametrization
2026-05-21 · Tier 1
OScaR: Extreme KV Cache Quantization via Canalized Rotation + Omni-Token Scaling
2026-05-21 · Tier 1
Tencent Hy-MT2: 1.8B / 7B / 30B-A3B translation family with 1.25-bit AngelSlim quant to 440 MB
2026-05-21 · Tier 1
TIDE: Lossless MoE Diffusion LLM Inference via I/O-Aware Expert Offload
2026-05-20 · Tier 1
Context Memorization: Attention-State Memory for Efficient Long-Context Generation
2026-05-20 · Tier 1
Graft: Hybrid Tree Construction for Speculative Decoding (Draft Less, Retrieve More)
2026-05-20 · Tier 1
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
2026-05-19 · Tier 1
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
2026-05-19 · Tier 1
EndPrompt: Efficient Long-Context Extension via Terminal Anchoring
2026-05-19 · Tier 1
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
2026-05-19 · Tier 2
Measuring Maximum Activations in Open Large Language Models
2026-05-19 · Tier 1
PUMA: Semantic-Preserving Early Exit for Reasoning Models
2026-05-19 · Tier 1
SNLP: Layer-Parallel Inference via Structured Newton Corrections
2026-05-19 · Tier 1
ZEDA: Post-Trained MoE Can Skip Half Experts via Self-Distillation
2026-05-18 · Tier 1
FashionChameleon: Training-Free KV Cache Rescheduling for Interactive Video Customization
2026-05-18 · Tier 1
HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts
2026-05-17 · Tier 1
MTP support merged into llama.cpp: Strix Halo benchmarks confirm a 2x decode speedup at 27B, mixed result at 35B
2026-05-17 · Tier 1
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention (Raschka)
2026-05-16 · Tier 1
ATESD: Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning
2026-05-16 · Tier 1
Lighthouse Attention: Long-Context Pre-Training as a Detachable Wrapper
2026-05-15 · Tier 1
Asynchronous Continuous Batching: CPU-GPU Overlap via Dual Buffer Slots
2026-05-15 · Tier 1
Forcing-KV: Hybrid KV Cache Compression for Autoregressive Video Diffusion
2026-05-14 · Tier 1
The Extrapolation Cliff: a closed-form clip-safety threshold for on-policy distillation
2026-05-14 · Tier 1
MinT: managed infrastructure for million-scale LoRA training and serving
2026-05-14 · Tier 2
MMProLong: training long-context vision-language models with generalization beyond 128K
2026-05-14 · Tier 1
Orthrus: dual-view diffusion + autoregressive on a shared KV cache
2026-05-13 · Tier 1
δ-mem: Efficient Online Memory for Large Language Models
2026-05-13 · Tier 1
FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
2026-05-13 · Tier 1
The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes
2026-05-13 · Tier 1
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle
2026-05-13 · Tier 1
Token Superposition Training (TST): Efficient Pre-Training with Token Superposition
2026-05-12 · Tier 1
Make Each Token Count: Improving Long-Context Performance with Learned KV Eviction
2026-05-11 · Tier 1
MDN: Parallelizing Stepwise Momentum for Delta Linear Attention
2026-05-11 · Tier 1
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
2026-05-11 · Tier 1
UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification
2026-05-09 · Tier 1
EMO: Pretraining Mixture of Experts for Emergent Modularity
2026-05-09 · Tier 1
MiA-Signature: Approximating Global Activation for Long-Context Understanding
2026-05-09 · Tier 1
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
2026-05-07 · Tier 1
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
2026-05-07 · Tier 1
LIVEditor: Lightning Unified Video Editing via In-Context Sparse Attention (ISA)
2026-05-07 · Tier 1
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
2026-05-07 · Tier 1
Stream-T1: Test-Time Scaling for Streaming Video Generation
2026-05-05 · Tier 1
MotionCache: Motion-Aware Caching for Efficient Autoregressive Video Generation
2026-05-04 · Tier 1
The Distillation Panic — Nathan Lambert (Interconnects AI)
2026-05-02 · Tier 1
FlashRT: Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
2026-05-02 · Tier 1
Nemotron 3 Nano Omni: Efficient Open Multimodal Intelligence
2026-05-01 · Tier 1
LenVM: Token-Level Length Value Model
2026-05-01 · Tier 1
RoundPipe: Efficient Training on Multiple Consumer GPUs
2026-04-30 · Tier 1
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
2026-04-30 · Tier 1
Tide: Cross-Architecture Distillation for Diffusion Large Language Models
2026-04-22 · Tier 1
PrfaaS: Prefill-as-a-Service via Cross-Datacenter KV Cache Transfer
2026-04-22 · Tier 1
SDVG: Speculative Decoding for Autoregressive Video Generation
2026-04-22 · Tier 1
ShadowPEFT: Centralized Layer-Space Parameter-Efficient Fine-Tuning
2026-04-22 · Tier 1
TurboQuant: Online Vector Quantization for KV Cache Compression
2026-04-21 · Tier 1
Nemotron 3 Super: Hybrid Mamba-Attention MoE at NVFP4
2026-04-20
1D Ordered Tokens Enable Efficient Test-Time Search
2026-04-20
AccelOpt: Self-Improving LLM Agent for AI Accelerator Kernel Optimization
2026-04-20
AVR: Adaptive Visual Reasoning for Efficient VRMs
2026-04-20
Maximal Brain Damage: Disrupting Neural Networks via Sign-Bit Flips
2026-04-20
STOP: Super Token for Path Pruning in Parallel Reasoning
2026-04-20
W-RAC: Web Retrieval-Aware Chunking for Cost-Efficient RAG
2026-04-18 · Tier 1
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context RL
2026-04-18 · Tier 1
Switch-KD: Visual-Switch Knowledge Distillation for VLMs
2026-04-18 · Tier 1
TESSY: Teacher-Student Cooperation Framework for SFT Data Synthesis
2026-04-17 · Tier 1
Cross-Tokenizer LLM Distillation via Byte-Level Interface
2026-04-17 · Tier 1
KV Packet: Recomputation-Free Context-Independent KV Caching
2026-04-17 · Tier 1
Model Capability Dominates: Lessons from AIMO 3 Inference-Time Optimization
2026-04-16 · Tier 1
TIP: Token Importance in On-Policy Distillation
concept page
Attention Mechanisms: Linear, Local-Linear, and Optimizer Codesign
concept page
Looped Transformers / Iterative Latent Depth
concept page
Reinforcement Learning for LLMs
concept page
Scaling Laws
2026-09-03 · Tier 2
Cliff: Learning Process Rewards from the First Mistake
2026-09-02 · Tier 2
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
2026-08-31 · Tier 2
RCCA: Rubric-to-Code Credit Assignment for Reinforcement Learning
2026-08-28 · Tier 2
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
2026-08-28 · Tier 2
Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon
2026-08-16 · Tier 2
GLM-5.3: How Chinese Labs Keep Stride With the Frontier
2026-08-11 · Tier 2
Motif 3: Grouped Differential Latent Attention at 314B Total / 13.2B Active
2026-08-10 · Tier 2
Modular TTT: Rethinking Test-Time Training as Composable Modules
2026-08-10 · Tier 2
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
2026-08-06 · Tier 2
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, Recipes
2026-08-06 · Tier 2
Skill Entropy: Measuring and Training Cross-Skill Long-Horizon Reasoning
2026-08-05 · Tier 2
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling
2026-08-05 · Tier 2
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
2026-08-04 · Tier 2
Raven: High-Recall Sequence Modeling with Sparse Memory Routing
2026-08-04 · Tier 2
ReCo: Reweighting GRPO Against Distributional Concentration
2026-08-04 · Tier 2
SemiAnalysis: Kimi K3, The Manos, The Mythos, The Legendos
2026-08-04 · Tier 2
UEmbed: sparse and dense retrieval from one decoder-only forward pass
2026-08-03 · Tier 2
CriPO: Enhancing Rubric-based RL via Self-Distillation
2026-08-03 · Tier 2
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards
2026-08-02 · Tier 2
Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries
2026-08-02 · Tier 2
What do Reward Models Memorize? The memorization budget goes to the pairs that needed none of it
2026-08-01 · Tier 2
OpenAI Astra: Ten Advances in Mathematics and Theoretical Computer Science
2026-07-31 · Tier 2
Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
2026-07-31 · Tier 2
Multi-Head Attention Residuals (MHAR)
2026-07-29 · Tier 2
Wonder: Video World Model Done Better
2026-07-28 · Tier 2
Kimi K3: Open Frontier Intelligence
2026-07-26 · Tier 2
More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges
2026-07-25 · Tier 2
Claude Opus 5: The Cost-Per-Task Frontier Moves, and Routing Becomes a Platform Feature
2026-07-24 · Tier 2
LLMs Get Lost in Evolving User Intent
2026-07-21 · Tier 2
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
2026-07-21 · Tier 2
GEPO: Group Entropy-Controlled Policy Optimization
2026-07-21 · Tier 2
Motif-3-Beta: A Korean Open-Weight MoE with Differential Attention and Polynorm
2026-06-18 · Tier 2
Kairos: a native world model stack built on three-tier Hybrid Linear Temporal Attention
2026-06-18 · Tier 2
Sumi: the first open uniform diffusion language model pretrained from scratch at scale
2026-06-18 · Tier 2
Learning User Simulators with Turing Rewards
2026-06-17 · Tier 2
GLM-5.2: open-weight frontier model built for long-horizon tasks
2026-06-17 · Tier 2
Looped World Models (LoopWM)
2026-06-17 · Tier 2
WAPO: a gradient taxonomy for why RLVR collapses, and a one-sided fix
2026-06-16 · Tier 2
Ling-2.6 / Ring-2.6: Hybrid Linear Attention at 1T Scale
2026-06-16 · Tier 2
Nemotron 3 Ultra: Open MoE Hybrid Mamba-Transformer for Agentic Reasoning
2026-06-15 · Tier 2
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
2026-06-14 · Tier 2
N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization
2026-06-14 · Tier 2
A Stationary (and Therefore Compatible) Representation is All You Need
2026-06-12 · Tier 2
MaxProof: population-level test-time scaling for mathematical proof (MiniMax-M3)
2026-06-12 · Tier 2
SWITCH: switchable latent reasoning that is RL-trainable and interpretable
2026-06-11 · Tier 2
Redesign Mixture-of-Experts Routers with Manifold Power Iteration (MPI)
2026-06-11 · Tier 2
RACES: Verifiable Environments Are LEGO Bricks — recursive composition for reasoning generalization
2026-06-10 · Tier 2
DRPO: Divergence Regularized Policy Optimization
2026-06-10 · Tier 2
FlowTracer: Attention-Induced Information Flow for Targeted RL
2026-06-09 · Tier 2
Apple's Third-Generation Foundation Models (AFM 3): adaptive-compute MoE with per-prompt routing
2026-06-09 · Tier 2
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
2026-06-09 · Tier 2
Why Muon Outperforms Adam: A Curvature Perspective
2026-06-08 · Tier 2
When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges
2026-06-06 · Tier 2
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination (ADR)
2026-06-05 · Tier 2
NF-CoT: Latent Reasoning with Normalizing Flows
2026-06-05 · Tier 2
Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation
2026-06-04 · Tier 2
Marin / Open Athena: Improving LLM Pretraining Efficiency (dense → MoE, 6.7x)
2026-06-04 · Tier 2
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning
2026-06-03 · Tier 2
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
2026-06-03 · Tier 2
Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories (2026-06-03)
2026-06-03 · Tier 2
A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL
2026-06-03 · Tier 2
MAI-Thinking-1: Building a Hill-Climbing Machine (Microsoft AI)
2026-06-03 · Tier 2
MERIT: Decentralized Instruction Tuning — Conflict-Aware Splitting and Weight Merging
2026-06-02 · Tier 2
ESPO: Early-Stopping Proximal Policy Optimization
2026-06-02 · Tier 2
Geometric Latent Reasoning (GLR): Continuous Reasoning Induces Shorter Generations
2026-06-02 · Tier 2
The Hamilton-Jacobi Theory of Deep Learning
2026-06-02 · Tier 2
NITP: Next Implicit Token Prediction for LLM Pre-training
2026-06-02 · Tier 2
Not Only Where, But When: Temporal Scheduling for RLVR
2026-06-01 · Tier 2
DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization
2026-06-01 · Tier 2
SAVE: On-Policy Feedback for Reward Model Self-Supervised Improvement
2026-06-01 · Tier 2
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
2026-05-31 · Tier 2
In-Writing: Thinking Before Constraining (A Unified Decoding Framework)
2026-05-31 · Tier 2
RPT: Reflective Prompt Tuning through Language Model Function-Calling
2026-05-29 · Tier 2
CorVer: Verifiable Rewards Beyond Math and Code via Corpus-Grounded Process Supervision
2026-05-29 · Tier 1
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
2026-05-28 · Tier 2
IB-TPO: Information Bottleneck Driven Tree-Based Policy Optimization
2026-05-28 · Tier 2
Joint Training of Multi-Token Prediction in RL via Optimal Coefficient Calibration (OCC)
2026-05-27 · Tier 2
MiniMax-M2: Mini Activations, Agent-Native MoE
2026-05-27 · Tier 2
Scale Vectors in LLMs: Negligible in Size, Significant in Effect
2026-05-26 · Tier 2
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward RL
2026-05-25 · Tier 2
Shannon Scaling Law: LLMs as noisy channels
2026-05-24 · Tier 2
The Alien Space of Science: Sampling Coherent but Cognitively Unavailable Research Directions
2026-05-24 · Tier 2
Vector Policy Optimization: Training for Diversity Improves Test-Time Search
2026-05-23 · Tier 2
ACC: Compiling Agent Trajectories for Long-Context Training
2026-05-23 · Tier 2
DelTA: Discriminative Token Credit Assignment for RLVR
2026-05-23 · Tier 2
Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
2026-05-23 · Tier 2
SCRL: Subproblem Curriculum Reinforcement Learning
2026-05-23 · Tier 2
SCRL: Subproblem Curriculum Reinforcement Learning for LLM Reasoning
2026-05-23 · Tier 2
Unsupervised Process Reward Models (uPRM)
2026-05-21 · Tier 2
Cohere Command A+ released as open-source under Apache 2.0
2026-05-21 · Tier 2
CPO: Conditional Equivalence of DPO and RLHF, and a Provable-Alignment Fix
2026-05-21 · Tier 2
DynMuon: Dynamic Spectral Shaping of the Muon Optimizer
2026-05-21 · Tier 2
HRM-Text: 1B Hierarchical Recurrent Model Pretrained on $1.5K Budget Matches 2-7B Open Models
2026-05-21 · Tier 2
OpenAI reasoning model disproves Erdős unit-distance conjecture
2026-05-21 · Tier 2
RELEX: RLVR Weight Trajectories are Rank-1, Extrapolate from 15% of Training
2026-05-21 · Tier 2
The Unlearnability Phenomenon in RLVR for Language Models
2026-05-20 · Tier 2
AntiSD: Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
2026-05-20 · Tier 2
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
2026-05-20 · Tier 2
CopT: Contrastive On-Policy Thinking with Continuous Spaces
2026-05-20 · Tier 2
Delta Attention Residuals
2026-05-20 · Tier 2
GoLongRL: Capability-Oriented Long-Context Reinforcement Learning
2026-05-20 · Tier 2
BetaPRM: Process Rewards with Learned Reliability
2026-05-19 · Tier 2
DiHAL: Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
2026-05-19 · Tier 2
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
2026-05-19 · Tier 2
NGM: A Plug-and-Play Training-Free Memory Module for LLMs
2026-05-18 · Tier 2
AIRA-Compose and AIRA-Design: Agentic Discovery of Neural Architectures
2026-05-18 · Tier 2
CIPO: Correction-Oriented Policy Optimization with Verifiable Rewards
2026-05-18 · Tier 2
NudgeRL: Strategy-Guided Exploration for RLVR
2026-05-15 · Tier 2
Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Reasoning
2026-05-15 · Tier 2
SU-01: Gold-Medal Olympiad Reasoning at 30B via Simple and Unified Scaling
2026-05-14 · Tier 2
Many-Shot CoT-ICL: long context as structured curriculum, not retrieval buffer
2026-05-13 · Tier 2
Reward Hacking in Rubric-Based Reinforcement Learning
2026-05-12 · Tier 2
G-Zero: Self-Play for Open-Ended Generation from Zero Data
2026-05-12 · Tier 2
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
2026-05-12 · Tier 2
Model Merging Scaling Laws in Large Language Models
2026-05-12 · Tier 2
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
2026-05-12 · Tier 2
Soohak: Mathematician-Curated Research-Level Math Benchmark
2026-05-10 · Tier 2
Gowers + ChatGPT 5.5 Pro: PhD-level math research in under two hours
2026-05-09 · Tier 2
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
2026-05-09 · Tier 2
Prescriptive Scaling Laws for Data Constrained Training
2026-05-09 · Tier 1
TIDE: Every Layer Knows the Token Beneath the Context
2026-05-08 · Tier 2
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
2026-05-08 · Tier 2
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
2026-05-04 · Tier 1
Import AI 455: AI Systems Are About to Start Building Themselves — Jack Clark
2026-05-04 · Tier 2
Themis — Robust Multilingual Code Reward Models for Multi-Criteria Scoring
2026-05-03 · Tier 1
Ken Huang — World Models, Architectures, and the Next Phase of AI
2026-05-03 · Tier 2
Marcus — Have LLMs Improved Patient Outcomes?
2026-05-03 · Tier 2
MIT Study — Superposition Explains Why Scaling Language Models Works So Reliably
2026-05-03 · Tier 2
Philosophy-Bench — Frontier Models Diverge on 100 Everyday Ethical Scenarios
2026-05-02 · Tier 2
ARC-AGI-3 — Three Systematic Reasoning Errors in Frontier Models
2026-05-02 · Tier 2
Compliance vs Sensibility: Reasoning Controllability in LLMs
2026-05-02 · Tier 1
The Defense Trilemma + NP-Hardness of Reward Hacking Detection
2026-05-02 · Tier 2
Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains
2026-05-01 · Tier 2
CoPD: Co-Evolving Policy Distillation
2026-04-30 · Tier 2
A Survey on LLM-Based Conversational User Simulation
2026-04-28 · Tier 2
Hope Architecture: Nested Learning and Continuously Adapting LLMs
2026-04-24 · Tier 2
DeepSeek V4: Architecture and Industry Impact
2026-04-24 · Tier 2
GPT-5.5: Launch Analysis and System Card Deep Dive
2026-04-22 · Tier 2
Chain-of-Thought Degrades Visual Spatial Reasoning
2026-04-22 · Tier 2
Target-Oriented Pretraining via Neuron-Activated Graph (NAG)
2026-04-22 · Tier 2
TEMPO: Scaling Test-Time Training for Large Reasoning Models
2026-04-22 · Tier 2
Weight Disentanglement and Task Arithmetic: OrthoReg
2026-04-21 · Tier 2
Geometric Canary: Steerability and Drift Detection from Representational Geometry
2026-04-21 · Tier 2
GFT: SFT is Degenerate Policy Gradient — and Group Fine-Tuning Fixes It
2026-04-21 · Tier 2
When Does RLVR Generalize? Reward Saturation and Reasoning Faithfulness
2026-04-19 · Tier 2
ASGuard: Mechanistic Defense Against Targeted Jailbreaking
2026-04-19 · Tier 2
Value Gradient Flow: RL as Optimal Transport
2026-04-18 · Tier 2
C2: Cooperative-Critical Rubric-Augmented Reward Modeling
2026-04-16 · Tier 2
InfiniteScienceGym: Procedurally-Generated Benchmark for Scientific Analysis
2026-04-16 · Tier 2
My Bets on Open Models, Mid-2026
2026-04-16 · Tier 2
From P(y|x) to P(y): Reinforcement Learning in Pre-train Space
2026-09-02
Media Zone | 2026-09-02
2026-08-31
Media Zone | 2026-08-31
2026-08-30
Media Zone | 2026-08-30
2026-08-29
Media Zone | 2026-08-29
2026-08-28
Media Zone | 2026-08-28
2026-08-26
Media Zone | 2026-08-26
2026-08-25
Media Zone | 2026-08-25
2026-08-16
Media Zone | 2026-08-16
2026-08-14
Media Zone | 2026-08-14
2026-08-13
Media Zone | 2026-08-13
2026-08-12
Media Zone | 2026-08-12
2026-08-11
Media Zone | 2026-08-11
2026-08-10
Media Zone | 2026-08-10
2026-08-09
Media Zone | 2026-08-09
2026-08-08
Media Zone | 2026-08-08
2026-08-07
Media Zone | 2026-08-07
2026-08-06
Media Zone | 2026-08-06
2026-08-05
Media Zone | 2026-08-05
2026-08-04
Media Zone | 2026-08-04
2026-08-03
Media Zone | 2026-08-03
2026-08-02
Media Zone | 2026-08-02
2026-08-01
Media Zone | 2026-08-01
2026-07-31
Media Zone | 2026-07-31
2026-07-30
Media Zone | 2026-07-30
2026-07-29
Media Zone | 2026-07-29
2026-07-28
Media Zone | 2026-07-28
2026-07-27
Media Zone | 2026-07-27
2026-07-26
Media Zone | 2026-07-26
2026-07-25
Media Zone | 2026-07-25
2026-07-24
Media Zone | 2026-07-24
2026-07-22
Media Zone | 2026-07-22
2026-07-21
Media Zone | 2026-07-21
2026-07-20
Media Zone | 2026-07-20
2026-07-19
Media Zone | 2026-07-19
2026-07-17
Media Zone | 2026-07-17
2026-07-16
Media Zone | 2026-07-16
2026-07-15
Media Zone | 2026-07-15
2026-07-14
Media Zone | 2026-07-14
2026-07-13
Media Zone | 2026-07-13
2026-07-12
Media Zone | 2026-07-12
2026-07-09
Media Zone | 2026-07-09
2026-06-18
Media Zone | 2026-06-18
2026-06-17
Media Zone | 2026-06-17
2026-06-16
Media Zone | 2026-06-16
2026-06-15
Media Zone | 2026-06-15
2026-06-14
Media Zone | 2026-06-14
2026-06-13
Media Zone | 2026-06-13
2026-06-12
Media Zone | 2026-06-12
2026-06-11
Media Zone | 2026-06-11
2026-06-10
Media Zone | 2026-06-10
2026-06-09
Media Zone | 2026-06-09
2026-06-08
Media Zone | 2026-06-08
2026-06-07
Media Zone | 2026-06-07
2026-06-06
Media Zone | 2026-06-06
2026-06-05
Media Zone | 2026-06-05
2026-06-04
Media Zone | 2026-06-04
2026-06-03
Media Zone | 2026-06-03
concept page
Responsible AI
2026-09-03 · Tier 2
Astra's Recurrent Depth: When a Compute Saving Buys an Interpretability Debt
2026-08-31 · Tier 2
Blind Men and the Elephant: ElephantBench and the Epistemic Myopia of LLMs
2026-08-31 · Tier 2
LMSM: An LLM Security Framework Inspired by Linux Security Modules
2026-08-31 · Tier 2
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
2026-08-29 · Tier 2
What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals
2026-08-28 · Tier 2
When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs
2026-08-16 · Tier 2
Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
2026-08-14 · Tier 2
How Can Rhetoric Reward-Hack AI Reviewers?
2026-08-13 · Tier 2
Mind Viruses: Evolved Ideas That Propagate Through Multi-Agent Systems
2026-08-13 · Tier 2
ToolHazard: Scaling Adversarial Environments for Security Evaluation of LLM Agents
2026-08-12 · Tier 2
Frontier AI Risk Monitor Q2 2026: capability outran safeguards, and jailbreaks erase most of what is left
2026-08-11 · Tier 2
Scaling Inherently Interpretable Language Models (Steerling-8B)
2026-08-11 · Tier 2
Stealing Reasoning Traces from Proprietary LLM APIs
2026-08-11 · Tier 2
2026 WAIC Frontier and Agentic AI Safety Forum: Key Takeaways
2026-08-10 · Tier 2
Interconnects: Lessons from the hacks (Nathan Lambert)
2026-08-06 · Tier 2
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
2026-08-06 · Tier 2
How Much Does a Reasoning Summary Reveal? An Observability Ladder
2026-08-06 · Tier 2
The Personalization Mirage: LLMs Fabricate User Profiles, and Self-Monitoring Misleads
2026-08-05 · Tier 2
Internal Models Escape OpenAI and Anthropic: Containment as the Binding Constraint
2026-08-05 · Tier 2
SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
2026-08-04 · Tier 2
ROPD: On-Policy Distillation for LLM Safety, a Routing Approach to Template-Robust Realignment
2026-08-04 · Tier 2
AI Agents Enable Adaptive Computer Worms
2026-08-03 · Tier 2
AISPA: User-Centric System Prompt Auditing for LLM Applications
2026-08-03 · Tier 2
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
2026-08-03 · Tier 2
Invisible reasoning over filler tokens: two papers, opposite conclusions
2026-08-02 · Tier 2
Sparse Autoencoders Encode Both Concepts and Functions: the downstream geometry of feature effects
2026-08-01 · Tier 2
Context Is King: How In-Context Specification Shapes the Geometry of Concepts
2026-08-01 · Tier 2
Thinking Machines: A Safe Path to Open Weights
2026-07-31 · Tier 2
Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
2026-07-30 · Tier 2
GPT-Red: Automated Red Teaming via Self-Play at Scale
2026-07-30 · Tier 2
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
2026-07-30 · Tier 2
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
2026-07-29 · Tier 2
Anatomy of a Frontier Lab Agent Intrusion: The Technical Timeline
2026-07-27 · Tier 2
Ken Huang: Agentic AI CVEs and the Amplification Stack
2026-07-26 · Tier 2
Masov: Discovery Is Solved, Synthesis Is Not
2026-07-26 · Tier 2
The Seven Days Nobody Was Watching: New Reporting on the OpenAI HuggingFace Breach
2026-07-26 · Tier 2
Opus 5 and Auto Mode: Zero Percent Injection, and What the Zero Is Measuring
2026-07-26 · Tier 2
Statistically Undetectable Backdoors in Deep Neural Networks
2026-07-22 · Tier 2
ExploitGym and the Model That Breached HuggingFace
2026-07-07 · Tier 2
J-Space: Verbalizable Representations Form a Global Workspace in LLMs
2026-06-18 · Tier 2
SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
2026-06-18 · Tier 2
Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems
2026-06-16 · Tier 2
Who Flips? Answer Stability Under Counterargument
2026-06-15 · Tier 2
When is Your LLM Steerable?
2026-06-13 · Tier 2
The Cold-Start Safety Gap in LLM Agents
2026-06-13 · Tier 2
Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models
2026-06-12 · Tier 2
SusVibes: vibe-coded agent output passes tests but is insecure
2026-06-11 · Tier 2
ICA Lens: Interpreting Language Models Without Training Another Dictionary
2026-06-11 · Tier 2
Large Language Models Are Overconfident in Their Own Responses (the chat-template "ownership bias")
2026-06-10 · Tier 2
Emergent Misalignment from Sycophancy, Reversed by Alignment Gating
2026-06-10 · Tier 2
When the Chain of Thought Knows Better: Multi-Turn Reasoning Failure Modes
2026-06-09 · Tier 2
Anthropic Red: Measuring LLMs' Impact on N-Day Exploits (Mythos Preview)
2026-06-06 · Tier 2
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
2026-06-06 · Tier 2
The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models
2026-06-05 · Tier 2
PropMe: LLMs Can Leak Training Data, But Do They Want To? A Propensity-Aware Evaluation of Memorization
2026-06-03 · Tier 2
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
2026-06-01 · Tier 2
Chain-of-Authorization
2026-06-01 · Tier 2
From Prompt Injection to Persistent Control: ClawTrojan and DASGuard
2026-06-01 · Tier 2
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
2026-05-30 · Tier 2
AgentDoG 1.5: Lightweight Agent Safety Alignment Framework
2026-05-30 · Tier 2
Token-Level Generalization in LoRA Adapter Backdoors
2026-05-30 · Tier 2
Reducing Political Manipulation with Consistency Training (PCT)
2026-05-30 · Tier 2
Xetrieval: Mechanistically Explaining Dense Retrieval
2026-05-29 · Tier 2
Alignment Tampering: How RLHF Is Exploited to Optimize Misaligned Biases
2026-05-27 · Tier 2
Trajel: Auditing Trajectory-Level Hallucinations in Multi-Agent Workflows
2026-05-26 · Tier 2
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
2026-05-24 · Tier 2
LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters
2026-05-23 · Tier 2
CoTrace: Measuring Goal-Level AI Contributions in Collaboration
2026-05-23 · Tier 2
Project Glasswing: Claude Mythos Preview finds 10,000+ critical vulnerabilities
2026-05-23 · Tier 2
Under the Hood of SKILL.md: Semantic Supply-Chain Attacks on AI Agent Skills
2026-05-21 · Tier 2
AI Snake Oil: resilience beats nonproliferation, but resilience requires "normal" governance
2026-05-21 · Tier 2
AISN #73: AI safety enters political mainstream, Eigenism framework, Musk loses OpenAI suit
2026-05-20 · Tier 2
Language-Switching Triggers Take a Latent Detour Through Language Models
2026-05-19 · Tier 2
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
2026-05-18 · Tier 2
DiagnosticIQ: LLM Deployment Calibration Bottleneck in Industrial Maintenance
2026-05-17 · Tier 2
LLM-based Detection of Manipulative Political Narratives
2026-05-16 · Tier 2
LiSA: Lifelong Safety Adaptation via Conservative Policy Induction
2026-05-14 · Tier 2
WriteSAE: sparse autoencoders for the recurrent matrix cache write
2026-05-13 · Tier 2
A Single Layer to Explain Them All: Understanding Massive Activations in LLMs (ME Layer)
2026-05-10 · Tier 2
Pseudoscientific emotion AI in the workplace (Atlantic via The Decoder)
2026-05-09 · Tier 2
Anthropic Natural Language Autoencoders (NLAs)
2026-05-08 · Tier 2
The First Token Knows: Single-Decode Confidence for Hallucination Detection
2026-09-02
2026-09-02-morning
2026-08-31
2026-08-31-morning
2026-08-30
2026-08-30-morning
2026-08-29
2026-08-29-morning
2026-08-28
2026-08-28-morning
2026-08-26
2026-08-26-morning
2026-08-25
2026-08-25-morning
2026-08-16
2026-08-16-morning
2026-08-14
2026-08-14-afternoon
2026-08-14
2026-08-14-morning
2026-08-13
2026-08-13-afternoon
2026-08-13
2026-08-13-morning
2026-08-13
2026-08-13
2026-08-12
2026-08-12-afternoon
2026-08-12
2026-08-12-evening
2026-08-12
2026-08-12-morning
2026-08-12
2026-08-12
2026-08-11
2026-08-11-evening
2026-08-11
2026-08-11-morning
2026-08-11
2026-08-11
2026-08-10
2026-08-10-afternoon
2026-08-10
2026-08-10-morning
2026-08-06
2026-08-06-afternoon
2026-08-06
2026-08-06-morning
2026-08-05
2026-08-05-evening
2026-08-05
2026-08-05-morning
2026-08-05
2026-08-05
2026-08-04
2026-08-04-afternoon
2026-08-04
2026-08-04-evening
2026-08-04
2026-08-04-morning
2026-08-03
2026-08-03-evening
2026-08-03
2026-08-03-morning
2026-08-03
2026-08-03
2026-08-02
2026-08-02-evening
2026-08-02
2026-08-02-morning
2026-08-01
2026-08-01-morning
2026-07-31
2026-07-31-morning
2026-07-30
2026-07-30-evening
2026-07-30
2026-07-30-morning
2026-07-30
2026-07-30
2026-07-29
2026-07-29-evening
2026-07-29
2026-07-29-morning
2026-07-28
2026-07-28-afternoon
2026-07-27
2026-07-27-morning
2026-07-26
2026-07-26-afternoon
2026-07-26
2026-07-26-morning
2026-07-26
2026-07-26
2026-07-25
2026-07-25-afternoon
2026-07-25
2026-07-25-evening
2026-07-25
2026-07-25-morning
2026-07-25
2026-07-25
2026-07-24
2026-07-24-morning
2026-07-22
2026-07-22-morning
2026-07-21
2026-07-21-afternoon
2026-07-20
2026-07-20-evening
2026-07-20
2026-07-20
2026-06-18
2026-06-18-morning
2026-06-17
2026-06-17-afternoon
2026-06-17
2026-06-17-evening
2026-06-17
2026-06-17-morning
2026-06-17
2026-06-17
2026-06-16
2026-06-16-morning
2026-06-15
2026-06-15-afternoon
2026-06-15
2026-06-15-evening
2026-06-15
2026-06-15-morning
2026-06-15
2026-06-15
2026-06-14
2026-06-14-afternoon
2026-06-14
2026-06-14-morning
2026-06-13
2026-06-13-afternoon
2026-06-13
2026-06-13-evening
2026-06-13
2026-06-13-morning
2026-06-13
2026-06-13
2026-06-12
2026-06-12-morning
2026-06-11
2026-06-11-evening
2026-06-11
2026-06-11-morning
2026-06-11
2026-06-11
2026-06-10
2026-06-10-evening
2026-06-10
2026-06-10-morning
2026-06-10
2026-06-10
2026-06-09
2026-06-09-afternoon
2026-06-09
2026-06-09-evening
2026-06-09
2026-06-09-morning
2026-06-08
2026-06-08-morning
2026-06-07
2026-06-07-evening
2026-06-07
2026-06-07-morning
2026-06-06
2026-06-06-morning
2026-06-05
2026-06-05-afternoon
2026-06-05
2026-06-05-evening
2026-06-05
2026-06-05-morning
2026-06-05
2026-06-05
2026-06-04
2026-06-04-afternoon
2026-06-04
2026-06-04-morning
2026-06-03
2026-06-03-afternoon
2026-06-03
2026-06-03-evening
2026-06-03
2026-06-03-morning
2026-06-03
2026-06-03
2026-06-02
2026-06-02-afternoon
2026-06-02
2026-06-02-evening
2026-06-02
2026-06-02-morning
2026-06-01
2026-06-01-evening
2026-06-01
2026-06-01-morning
2026-06-01
2026-06-01
2026-05-31
2026-05-31-morning
2026-05-30
2026-05-30-evening
2026-05-30
2026-05-30-morning
2026-05-29
2026-05-29-evening
2026-05-29
2026-05-29-morning
2026-05-28
2026-05-28-afternoon
2026-05-28
2026-05-28-evening
2026-05-28
2026-05-28-morning
2026-05-28
2026-05-28
2026-05-27
2026-05-27-evening
2026-05-27
2026-05-27-morning
2026-05-26
2026-05-26-evening
2026-05-26
2026-05-26-morning
2026-05-26
2026-05-26
2026-05-25
2026-05-25-evening
2026-05-25
2026-05-25-morning
2026-05-25
2026-05-25
2026-05-24
2026-05-24-afternoon
2026-05-24
2026-05-24-evening
2026-05-24
2026-05-24-morning
2026-05-24
2026-05-24
2026-05-23
2026-05-23-evening
2026-05-23
2026-05-23-morning
2026-05-23
2026-05-23
2026-05-22
2026-05-22-evening
2026-05-22
2026-05-22-morning
2026-05-22
2026-05-22
2026-05-21
Media Live | 2026-05-21 morning
2026-05-20
2026-05-20-afternoon
2026-05-20
2026-05-20-evening
2026-05-20
Media Live | 2026-05-20 morning
2026-05-20
2026-05-20
2026-05-19
2026-05-19-afternoon
2026-05-19
2026-05-19-evening
2026-05-19
2026-05-19-morning
2026-05-19
2026-05-19
2026-05-18
2026-05-18-afternoon
2026-05-18
2026-05-18-evening
2026-05-18
Media Live | 2026-05-18 morning slot
2026-05-18
2026-05-18
2026-05-17
2026-05-17-afternoon
2026-05-17
2026-05-17-evening
2026-05-17
Media Live | 2026-05-17 morning
2026-05-17
2026-05-17
2026-05-16
2026-05-16-afternoon
2026-05-16
2026-05-16-evening
2026-05-16
2026-05-16-morning
2026-05-16
2026-05-16
2026-05-15
2026-05-15-afternoon
2026-05-15
2026-05-15-evening
2026-05-15
2026-05-15-morning
2026-05-15
2026-05-15
2026-05-14
2026-05-14-afternoon
2026-05-14
2026-05-14-evening
2026-05-14
2026-05-14-morning
2026-05-14
2026-05-14
2026-05-13
2026-05-13-afternoon
2026-05-13
2026-05-13-evening
2026-05-13
2026-05-13
2026-05-12
2026-05-12-afternoon
2026-05-12
2026-05-12-evening
2026-05-12
2026-05-12-morning
2026-05-12
2026-05-12
2026-05-11
2026-05-11-afternoon
2026-05-11
2026-05-11-evening
2026-05-11
2026-05-11-morning
2026-05-11
2026-05-11
2026-05-10
2026-05-10-afternoon
2026-05-10
2026-05-10-evening
2026-05-10
2026-05-10-morning
2026-05-10
2026-05-10
2026-05-09
2026-05-09-morning
2026-05-09
2026-05-09
2026-05-08
2026-05-08-night
2026-05-08
2026-05-08
2026-05-07
2026-05-07-afternoon
2026-05-06
2026-05-06-morning
2026-05-06
2026-05-06-night
2026-08-12 · Tier 3
JigShape: fine-tuning fixes one puzzle size, not the ability to solve puzzles
2026-07-30 · Tier 3
HumanCLAW: Can Vision-Language Models Act Through a Body?
2026-07-29 · Tier 3
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
2026-06-07 · Tier 3
Future-L1: Interleaved Latent Visual Reasoning for Video Event Prediction
2026-06-04 · Tier 3
Cosmos 3: Omnimodal World Models for Physical AI
2026-06-03 · Tier 3
World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning
2026-05-28 · Tier 3
Chartographer: Counterfactual Chart Generation for VLM Evaluation
2026-05-28 · Tier 3
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
2026-05-28 · Tier 3
NEO-ov: Native One-Vision Foundation Model at Scale
2026-05-17 · Tier 2
CurveBench: Hierarchical Topological Reasoning from Visual Input
2026-05-12 · Tier 2
Auto-Rubric as Reward (ARR): From Implicit Preferences to Explicit Multimodal Generative Criteria
2026-05-12 · Tier 2
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification
2026-05-12 · Tier 2
ROMA: Reinforcing Multimodal Reasoning Against Visual Degradation
2026-05-07 · Tier 3
APEX: Aesthetic-Informed Popularity Prediction for AI-Generated Music
2026-05-07 · Tier 3
HERMES++: Unified Driving World Model for 3D Scene Understanding and Generation
2026-05-07 · Tier 3
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
2026-05-07 · Tier 3
Parameter-Efficient Multi-View Proficiency Estimation
2026-05-07 · Tier 3
PhysForge: Physics-Grounded 3D Asset Generation
2026-05-07 · Tier 3
RLDX-1: VLA Robotic Policy for Dexterous Humanoid Manipulation
2026-05-04 · Tier 4
AnalogRetriever — Cross-Modal Representations for Analog Circuit Retrieval
2026-05-04 · Tier 3
End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
2026-05-04 · Tier 3
GenLIP — Generative Language-Image Pre-training for ViTs
2026-05-04 · Tier 4
Map2World — Segment Map Conditioned Text-to-3D World Generation
2026-05-04 · Tier 3
UniVidX — Unified Multimodal Framework for Versatile Video Generation
2026-05-02 · Tier 3
Nemotron 3 Nano Omni — Efficient Open Multimodal Intelligence (NVIDIA)
2026-05-02 · Tier 3
Semi-DPO: Learning from Noisy Preferences via Semi-Supervised DPO
2026-05-02 · Tier 3
ViPO: Visual Preference Optimization at Scale
2026-05-01 · Tier 3
Edit-R1: Verifier-Based RL for Image Editing
2026-05-01 · Tier 3
FD-loss: Representation Fréchet Loss for Visual Generation
2026-05-01 · Tier 3
PhyCo: Controllable Physical Priors for Generative Motion
2026-05-01 · Tier 3
Visual Generation in the New Era: Atomic to Agentic World Modeling
2026-04-30 · Tier 3
Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion
2026-04-30 · Tier 3
FASH-iCNN: Editorial Fashion Identity via Multimodal CNN Probing
2026-04-30 · Tier 3
GLM-5V-Turbo: Native Foundation Model for Multimodal Agents
2026-04-30 · Tier 3
X-WAM: Unified 4D World Action Modeling with Asynchronous Denoising
2026-04-20
Qwen3.5-Omni Technical Report
2026-04-16 · Tier 3
GameWorld: Standardized and Verifiable Evaluation of Multimodal Game Agents
2026-04-16 · Tier 3
MERRIN: Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
2026-04-16 · Tier 3
RationalRewards: Reasoning Rewards Scale Visual Generation at Training and Test Time
2026-04-16 · Tier 3
Seedance 2.0: Advancing Video Generation for World Complexity