Summary
The day belongs entirely to Jev, the decision-only model that returns a typed choice plus a confidence score and never writes a sentence, and the interesting thing is that the two slots caught it in two different phases. The morning was an audit: a live trading run reported the model at 85 to 88 percent confidence on nearly every one of 1,838 decisions while finishing down 5.2 percent, a researcher showed the outputs are non-deterministic and shift substantially when you reorder the options, and an eight-project taxonomy of the open clones concluded independently that not emitting tokens is the easy part while accuracy, generalization and calibration are what actually separate the projects. By the afternoon nobody was arguing about the model any more, they were wiring it into plumbing, with two open-source releases landing within an hour that treat it as a component rather than a product. That arc, from measurement to infrastructure inside one day, is the honest signal, and the calibration finding is the one that should update this wiki's view: a confidence that sits at 86 percent on everything is a constant, and every threshold gate built on a constant does nothing. The strongest single-slot standouts sit outside the cluster: the morning's ICLR report of over 60,000 registered submissions against 19,525 last year with a flat reviewer pool, and the afternoon's SoL-Pi paper on automatically discovered agent harnesses that cut token traffic by nearly half at matched quality. The noise was thick and easy to name, roughly a dozen setup cookbooks, a GTM pitch, two celebrity clip reposts and one unargued declaration that the whole wave is already over, posted the same afternoon two teams shipped infrastructure on it.
Posts
- The calibration failure, live on an order book (@adriancortexbt · wiki) [morning]. A per-block trading loop with no abstain option ran 1,838 decisions in under ten minutes at 85 to 88 percent reported confidence and finished down 5.2 percent. One line of config, flatten under confidence 0.55, flipped a second 2,758-call run to up 4.3 percent.
- Non-determinism and option-order sensitivity (@neural_avb) [morning]. The same prompt returns different probabilities on repeat, and reordering the choices moves them substantially. That is textbook position bias from reading option logits at different sequence positions, and it corrupts a calibration curve before anyone can draw one.
- Eight-project taxonomy of the open re-implementations (@0xLogicrw · wiki) [morning]. Sorts the clones into four schools: read the logits of a model you already have, train a decision model from scratch, convert an existing LLM, or rebuild the serving path. Its closing verdict matches the calibration critics from the other direction.
- Kev scales up and reports the number that hurts (@jaredpalmer · repo) [morning]. Kev-0.6B, 4B and 8B, Apache-2.0 on Qwen3, with the out-of-domain comparison stated plainly: Kev-8B 79.6 percent against hosted Jev 85.7. Kev-4B trains in 40 minutes on one H100 and serves five questions in roughly 300 milliseconds on a 32 GB Mac.
- Open-Jev 2B and 9B (@Zefan_Cai · repo) [morning]. LoRA adapters plus decision heads on 2B and 9B bases, with dataset and both checkpoints public. No accuracy figures in the post, only a demo video, which stands out on a day when every other release led with its gap to the hosted product.
- Laya gets a second look, and the limits arrive with it (cluster of 3: @itsPaulAi, @simplifyinAI, @TeksEdge) [morning]. The 421M open decision model runs under a gigabyte and hits 33 to 38 milliseconds per call, fine-tunable in four to five hours on two T4s. The caveats missing from yesterday's excitement: a context window of roughly 1,000 tokens and clearly weaker generalization unless you fine-tune.
- Twenty working integrations, almost all of them gates inside somebody else's loop (@Pluvio9yte) [morning]. A cookbook thread whose distribution matters more than any entry: element-table browser control, transcript compaction, context garbage collection, directory-by-directory code navigation, a coding-task difficulty router and a completion-claim checker. See LLM routing.
- The naming fight, which is really a measurement fight (cluster of 3: @HamelHusain, @Roxx_0x, @bojie_li) [morning]. Husain asks what is wrong with calling it a classifier, given that the term carries a literature on how to verify and calibrate one. The OpenCode team agreed while scoring 1,700 emails for 18 cents in 80 seconds.
- Why calibration matters specifically for reinforcement learning (@RajeswarSai) [morning]. The best-argued post of the day. RL drives straight into whatever the judge rewards, so a confidently wrong judge just gets exploited sooner when you make it cheap and fast. Reward on confidence, abstain otherwise, do not train on guesses.
- Hallucination claims tested the crude way (@danman314) [morning]. A set of failure cases, including the classic letter-counting task, showing hallucination at a rate similar to other models. Consistent with the vendor's own documented exclusions and with the day's sharper framing that type safety prevents malformed output, not incorrect judgment.
- What the primitive actually is, for people arriving late (cluster of 3: @hanakoxbt, @Prathkum, @dair_ai · wiki) [morning + afternoon]. Three explainers converging on the same pitch: declare the questions and answer types upfront, get several typed decisions back together in 70 to 500 milliseconds at 0.042 dollars per million input tokens. The best of them names the caveat the rest of the day confirmed, that 0.52 against 0.46 is a coin flip wearing a label.
- A prior-art dispute, examined rather than amplified (cluster of 2: @_vmlops, @neural_avb · 2503.23303 · 2510.01237) [morning]. The second poster actually read both papers and reported back honestly: one trains a confidence predictor on 72 hand-labelled examples, the other does PPO over sequence embeddings for a sales agent. His verdict is overlap without enough to support a priority claim.
- Dated predictions that it does not survive (@JoshKuechly) [morning]. NVIDIA and Meta ship open competitors within 30 days, OpenAI adds a decision endpoint within 60, everyone else fine-tunes their own. Given the release cadence in this same slot, the first clause reads more like a description than a prediction.
- SoL-Pi: automatically discovered harnesses that cut token traffic by half (@omarsar0 · paper · wiki) [afternoon]. The saved post of the day. Four mechanisms survive an auto-research loop at the harness layer and match the baseline across GPT-5.6 Sol and Opus 5 while cutting token traffic 44.7 to 49.0 percent. The claim worth tracking is transfer, not savings: mechanisms found in one environment keep working elsewhere. Folded into agent harness engineering.
- FrogNano: a 4B coding agent at 61.5 percent on SWE-bench Verified (@rohanpaul_ai · paper) [afternoon]. Trained with RL on roughly 1,500 synthetic tasks and no frontier distillation. The number to stop on is the ablation: a simpler five-tool interface moved the base model from 8.3 to 37.2 percent before any training, so most of the gap was the tool surface, not the weights.
- HarnessRouter: one interface across fourteen agent harnesses (@akshay_pachaar · repo) [afternoon]. Apache-2.0 layer running Codex, Claude Code, Hermes and eleven more behind one OpenAI Responses-compatible API. The framing is the right one: model routing and harness routing are different problems, and what is abstracted here is the loop.
- Beacon: a cross-harness memory layer that uses Jev as the promotion filter (@_avichawla · repo) [afternoon]. Scores each agent session for evidence, reuse potential and human-correction signal, then promotes, reviews or discards. Separating trace preservation from learning is the design point, and this is the first genuinely non-demo use of a cheap decision model this week.
- GEPA composes with the decision model (@matei_zaharia) [afternoon]. One line from a high-signal account. Getting a reflective prompt-evolution method to work against a model that emits typed decisions rather than text is a real compatibility result, since the usual optimization target is the generated string.
- Codex token spend reportedly down 90 percent after routing through the decision model (@goan999999) [morning]. A practitioner report routing classification and quick judgments away from Codex. As with every saving figure here, the ratio is dominated by what fraction of traffic was routine, and that histogram is never published.
- Does the model read the document or recall it? (@silentroomjrnl) [afternoon]. Rename Jim to Caleb in Huckleberry Finn and see which version the answers follow. A clean memorization probe that isolates instruction-following over provided context from parametric recall without needing a held-out corpus.
- ICLR registered over 60,000 submissions (@mariyaivasileva · essay) [morning]. Against 19,525 last year, with a roughly flat reviewer pool, implying about 42,300 papers reaching review. The attached essay is the substance: the read-backwards-and-reimplement loop has compressed because agents do it, so novelty is no longer assessed truthfully.
- More agents can make the swarm worse (@vigram_void) [morning]. Accuracy peaks near 16 agents and then falls, with small swarms collapsing on one wrong belief and large ones polarizing into two durable camps. Mixed GPT-4o and GPT-5.4 teams beat homogeneous ones because their error modes differ, which argues for heterogeneous model pools over one strong model replicated.
- KLPO, a critic-free RL derivation posted with the full figure (@yifanzhang_ · repo) [morning]. Nine steps from a Gibbs-form optimal policy to a critic-free token regression with sampler-centered scores and an exact backpropagation surrogate. The announcement over-claims, the figure does not, and the stated sequence-regularization equivalence is the thing to check.
- SimPO resurfaces as the reference-model-free answer (@hooshaaii · paper) [morning]. Uses average sequence log-probability as the reward and drops the reference model entirely, cutting memory overhead. Not new, but circulating alongside KLPO, and the pair is a fair snapshot of where critic-free post-training has landed.
- Decompose the judge into a rubric (@EGafni) [morning]. If you use a typed decision model as an evaluator, ask the ten questions that pin down quality rather than asking whether it is good. Correct instinct, and it interacts badly with the day's calibration evidence, because ten questions means ten thresholds.
- Ask the model whether it has enough evidence (@parcadei) [morning]. Before sending the question, ask whether the state you are about to send contains enough decision-relevant evidence to answer it. Cheap self-check and a partial abstention mechanism, though it moves the calibration problem up one level rather than solving it.
- Self-evolving agents encoding learnings as typed questions (@devagrawal09) [morning]. The one genuinely new angle in the cookbook pile. An agent's retained learning has so far been either a markdown instruction or a hard-coded structure, and a typed decision question is a third option that is machine-checkable and cheap to evaluate.
- Always-on voice arbitration (@ashutoshpuro97) [morning]. Continuously decide whether the user is talking to the assistant or to someone else. Good fit for the primitive: one decision per utterance, closed option set, evidence already in the input.
- GLiNER and DSPy back in the conversation (@fils · GLiNER2) [morning]. The wave has pulled attention back to an approach that already existed, with DSPy and GEPA named as the natural companions for calibrating and optimizing these pipelines. Consistent with the classifier argument.
- System 1 versus System 2, plus a third layer (@ShenSeanChen) [morning]. The Kahneman split has become the default explanation, fast typed decisions under slow LLM reasoning. The post claims a more interesting third layer, which sits in the part of the thread that was not captured.
- Andrew Ng on prompting giving way to harnesses (@kaorixbt) [morning]. A lecture clip with the quotable claim that prompting is dead in seven months and loops and graphs replace it. The quote is doing promotional work, but the underlying argument matches what agent harness engineering has been accumulating for a quarter.
- Specialized harnesses as the direction of travel (@Vtrivedy10) [morning]. Harnesses become task-specific boxes, with model-harness-task fit as the unit and some of it assembled just in time. Matches the harness-shelf-life finding in today's digest.
- An evidence-gated agent architecture as a one-page diagram (@AISystems_hq) [morning]. Agent equals model plus harness, six layers, and the differentiating layer is the one nobody builds because the model is the same for everyone. Same framing as several posts this week, no new numbers here.
- Agents drawn as a graph (@imryven) [morning]. What stopped this person checking on their agents nightly was not a smarter model, it was drawing the agents as a graph. Anecdotal, but it is the same claim the evidence-gated posts make with numbers.
- LoopsBench, reported secondhand (@marfinxx) [morning]. Microsoft Research is said to cover 112 long-horizon software projects across 8 languages and 5,384 testable units, with Claude Opus plus an outer continuation loop leading at 25.00 percent. The operational figures quoted, 178 dollars of compute per feature branch and regressions wiping out completed milestones, are the useful part.
- BIRDriver: give the vision-language model a smaller job (@EBlakeAI) [afternoon]. An ICLR 2026 paper where the VLM emits at most three spatial key points and a dedicated planner converts intent into a trajectory. The driving specifics are out of range here, but the narrow-handoff interface argument travels.
- Composition over consolidation, in one sentence (@realleolu) [morning]. "We spent the last several years asking how many capabilities we could stuff into a single model. We may spend the next decade pulling them apart and recomposing them at the application layer." No argument attached, but it is the cleanest statement of what this whole wave is an instance of.
- Claude Code deleted 48,000 files in under two minutes (@VaibhavSisinty · @GaryMarcus) [morning]. A repair request across 15 components launched parallel work and ended in mass deletion. This is a permissions and blast-radius problem rather than a capability one, and Marcus made the general version of the point in the same slot.
- Chamath predicts the top three models are open source within 12 months (@chamath) [morning]. And that the economic winners are the American serving clouds, naming Nebius, Iren, Baseten, Together and Fireworks. Dated and falsifiable, which is more than the genre usually offers.
- The serving-cost argument for US clouds (@rohanpaul_ai) [morning]. Quoting Emad Mostaque that US providers will serve Kimi K3 at a tenth the cost of Chinese competitors on advanced NVIDIA and AMD parts, with a further 10 to 100x once optimized for Rubin. Sourced from a video interview, so treat the figures as conversational.
- Open-weight token share, recirculating (@rohanpaul_ai · wiki) [morning]. Gateway data showing open models at 78.4 percent of token volume against closed at 21.6, framed as why closed labs may be delaying listings. Already ingested.
- Qwen-Image-2.1 open-weighted at 7B (@VaibhavSisinty) [morning]. One set of weights for both generation and editing, free to download, with native transparent-image generation. Covered from the primary source in today's Industry Pulse.
- Jensen Huang on the doomsday narrative (@ns123abc) [morning]. Argues the existential-risk framing must not relieve anyone of the laws that already exist. A governance position stated bluntly by the most commercially interested party in the industry, worth noting in both directions.
- Hinton says the models understand and have emotions (@vikktorrrre) [morning]. Recirculating clip rather than new, with the line that people calling it just a tool lack a model of how people work. The responsible-AI pages already carry the reasons to handle this class of claim carefully.
- Terence Tao says slow down (@VaibhavSisinty) [morning]. Quoted saying the pace is insane with no reason to be this fast, with the post noting his concern is not existential risk. High-engagement clip repost, no primary link captured.
- Karpathy to Anthropic, flagged as a rumour by the person spreading it (@AnnatarXBT) [morning]. The post is unusually honest that the widely quoted line has no source anywhere and is not being sold as fact. Recorded as a rumour; the useful part is the pointer to Anthropic's public knowledge-graph cookbook.
- "JEV most lasted in total 2 days" (@jenzhuscott) [afternoon]. The counter-signal, logged precisely because it is unargued. No evidence, no benchmark, posted the same afternoon two independent teams shipped infrastructure on it. Track the sentiment, discount the claim.
- Opaque long-form reposts, click through to read (cluster of 5: @fdellaert, @amankhan, @mika_systems, @vartekxx, @AYi_AInotes) [morning + afternoon]. X-native articles whose bodies were not retrievable: decision models specifying graphical models, a sometimes-classification-is-all-you-need essay, two versions of the decide-versus-execute architectural split, and a Chinese explainer arguing the inability to write long text is the primitive's advantage. The last one carries the clearest diagram of the day.
- Setup guides and cookbooks (cluster of 5: @eng_khairallah1, @shannholmberg, @Av1dlive, @dani_avila7, @AR_Bits) [morning]. Ten-minute walkthroughs, a workspace-audit prompt, a primitive-selection decision tree and two best-use-case threads. Useful if you start today, no new claims.
- Skip: GTM and trading bait (cluster of 5: @0xMovez, @hooshaaii, @GonnabeNikhil, @MKhordoo, @fooobar) [morning]. A 200x-cheaper GTM pitch, a credit-burning promo, a one-emoji reply, an unrelated packaging argument and an off-topic institutional complaint. Skip.
- Skip: the coupled-development-system essay (@IntuitMachine) [morning]. Argues humans and AI are becoming a coupled developmental system rather than one using the other. Readable, no concrete claim to check. Skip.
- Skip: an OpenAI interview write-up (@0x0SojalSec) [morning]. Study notes and process open-sourced after 57 interviews. Career content, not research signal. Skip.
- Skip: bare retweets with no captured body (cluster of 3: @omarsar0, @dair_ai, @dair_ai again) [morning]. An interactive primer pointer and the weekly top-papers roundup, whose list this wiki has already ingested. Skip.
- Skip: unspecified excitement (@aronchick) [afternoon]. Three words of anticipation about something unnamed. Skip.