Summary
The decision-model story did not slow down this afternoon, but its character changed: where the morning was people auditing whether Jev's confidence numbers mean anything, the afternoon is people wiring it into plumbing and explaining it to newcomers. Seven of the twelve posts touch Jev, and almost none of them argue about the model itself. Two open-source infrastructure releases landed within an hour, Beacon and HarnessRouter, both of which treat Jev as a component rather than a product, which is what adoption actually looks like. Against that, the single most substantive item is the saved post, an NVIDIA and NTU and MIT paper on discovering agent harnesses automatically that cuts token traffic by nearly half at matched quality, and the best unsaved item is a 4B coding agent reporting 61.5 percent on SWE-bench Verified with a number buried in it that deserves more attention than the headline: simplifying the tool interface alone took the base model from 8.3 to 37.2 percent. The noise is easy to name. One hype post with no content, and one drive-by declaring the whole Jev wave already over, which reads as sentiment rather than evidence.
Posts
SoL-Pi: automatically discovered harnesses that cut token traffic by half (@omarsar0 · paper). The saved post of the slot, and it lands on the reader's central question of where agent cost actually goes. Rather than hand-tuning a harness, the authors from NVIDIA, NTU and MIT run auto-research loops at the harness layer across many repository-derived and verifier-driven environments and keep only the mechanisms that survive selection. Four survive and form SoL-Pi: Action Fusion for how actions execute, Online Context Compact for compaction during a run, ObservationPack for observation handling, and an Evidence-Preserving Reducer for delegated reading. On the 51-task EdgeBench evaluation it matches the Pi baseline across GPT-5.6 Sol and Opus 5 while cutting recorded token traffic 44.7 to 49.0 percent and API cost by about a third, which they report as 8.75 to 13.50 dollars saved per hour against native Codex and Claude Code harnesses. The claim worth tracking is not the savings but the transfer: mechanisms found in one environment keep working elsewhere, which is what would make harness search worth running once rather than per-project. Already written up at SoL-Pi recursive harness research loops and folded into agent harness engineering.
FrogNano: a 4B coding agent at 61.5 percent on SWE-bench Verified, no frontier distillation (@rohanpaul_ai · paper). Starts from Qwen3.5-4B and trains with reinforcement learning on roughly 1,500 synthetic software-engineering tasks, with the curriculum regenerating as the model improves so problems stay hard but still learnable. The headline is the 61.5 percent after five rounds without distilling from a frontier model, but the number that should stop you is the ablation: switching to a simpler five-tool interface moved the base model from 8.3 to 37.2 percent before any training happened. Most of the gap between a small model and a useful agent was the tool surface, not the weights, which is the same finding the harness literature keeps arriving at from the other direction.
HarnessRouter: one interface across fourteen agent harnesses (@akshay_pachaar · repo). An Apache-2.0 infrastructure layer that runs Codex, Claude Code, Hermes, DeepSeek Harness, System One and nine more behind a single OpenAI Responses-compatible API, handling sessions, streaming, files, cancellation and failure uniformly. The harnesses run locally and the common task interface is specified as an open standard, the Unified Harness Protocol. The framing in the post is the useful part and is exactly right: model routing and harness routing are different problems, and the thing being abstracted here is the loop, not the model. Relevant to LLM routing as the layer above it.
Beacon: a cross-harness memory layer that uses Jev to decide what was worth learning (@_avichawla · repo). Captures full agent session history across Claude Code, Codex, Cursor, OpenCode and twenty-plus others, then scores each run for evidence, reuse potential and human-correction signal, and applies a policy that promotes, reviews or discards it. Promoted runs become reusable skills available to other agents on the project. The design point worth keeping is the separation the post makes explicitly: preserving a trace and learning from it are different operations, and most sessions are routine exploration and one-off fixes that should stay inspectable without becoming guidance. Using a cheap decision model as the filter on that promotion step is the first genuinely non-demo use of Jev this week.
What Jev actually is, for people arriving late (cluster of 2: @Prathkum, @dair_ai). Both are onboarding explainers now that access is open, and the first carries the concrete numbers: it reads text or JSON but never writes prose, returning typed decisions with a confidence on each, several per call, in roughly 70 to 500 milliseconds, at 0.042 dollars per million input tokens with output free. The framing offered is an AI you call like a function rather than one you chat with. Read this against the calibration reckoning from this morning, which found the reported confidences sitting at 85 to 88 percent on nearly everything: the pitch here is that your code acts when the model is sure and escalates when it is not, and that pitch is only as good as the calibration the morning posts were unable to confirm.
GEPA composes with Jev (@matei_zaharia). One line, but from someone whose signal-to-noise is high. GEPA is the reflective prompt-evolution method that optimizes a system by mutating and selecting textual components, and getting it to work against a model that emits typed decisions rather than text is a real compatibility result, since the usual optimization target is the generated string. Worth watching whether anyone posts numbers.
Jev plus graphical models, click through to read (@fdellaert). A long-form X article asking whether a decision-only model can specify graphical models and change how systems reason under uncertainty. The thesis is intriguing and the author is a serious researcher in the area, but the argument lives in the article body rather than the post. Click through to read.
Does the model read the document or recall it? (@silentroomjrnl). A clean memorization probe: take Huckleberry Finn, which every model has seen, rename Jim to Caleb, and see whether answers follow the text you supplied or the text in the weights. This is the right shape of test and it generalizes well beyond the one book, since it isolates instruction-following over provided context from parametric recall without needing a held-out corpus. The post reached far wider than anything else in the slot.
BIRDriver: give the vision-language model a smaller job (@EBlakeAI). An ICLR 2026 paper where the VLM does not drive. It reads the scene and emits at most three spatial key points, and a dedicated motion planner converts that intent into a trajectory. The driving specifics are outside this wiki's usual range, but the interface argument travels: put the semantic model where semantics are needed, keep geometry and control in the component that is good at them, and make the handoff narrow enough to inspect.
"JEV most lasted in total 2 days" (@jenzhuscott). The counter-signal, and worth logging precisely because it is unargued. No evidence, no benchmark, just a verdict that the wave is over, posted on the same afternoon two independent teams shipped infrastructure built on it. Track the sentiment, discount the claim. The three-day ecosystem census on the open clones is the better basis for judging whether this holds.
Unspecified excitement (@aronchick). Three words of anticipation about something unnamed. Skip.