Summary
The Jev story reached its measurement phase, and that is the whole slot: 22 of 50 posts are about the decision-only model, and the conversation has finally moved off hype onto numbers. Three things changed today. JevBench landed as the first independent leaderboard for the model class, with hosted Jev on top at 75.3 and the open SemIf a hair behind at 74.6, which quietly resolves the prediction this wiki filed on 09-16 that Jev needed an independent benchmark inside 30 days or should be treated as vapor. The first real skepticism arrived too, aimed squarely at yesterday's clone euphoria: Laya, the open model that looked like it beat Jev, turns out to cap at 512 to 1K context and score near coin-flip accuracy without fine-tuning first, which is a direct challenge to the commoditization claim this wiki made on 09-19. And an actual production cost teardown showed up, the sharpest number of the day: one engineer tagged all 1,284 tool calls in a single agent session, found 938 of them were binary yes-or-no questions, and cut a $64.77 bill to $26.16 by routing those to a classifier at $0.000041 and 38ms per decision instead of 1.84s. Outside Jev the signal is thin but real: Noam Brown's 80-minute Dwarkesh interview put concrete figures on OpenAI's Navier-Stokes result (10,000 agents, 130B tokens, 88 hours) while pouring cold water on multi-agent as the cause, a CUDA performance guide circulated well, and three separate posts pushed the argument that the AI safety lobby is really a regulatory-capture play. The tail of the slot is mostly grift and ads, including a trading bot that lost its builder $31,680 and got the highest engagement of the day for it, which is a good reminder that raw engagement is the worst possible ranking signal here.
Posts
- JevBench, the first independent leaderboard for decision-only models (@airesearch12 · leaderboard). Hosted Jev leads at 75.3, the open SemIf is second at 74.6, a gap of under one point. This is the measurement the 09-16 ingest asked for, and it says the capability replicates but the hosted version still edges the field. See Jev: a decision-only model that never writes a sentence.
- The openjev census: at least 18 re-implementations now tracked (cluster of 2: @airesearch12 · @KUMAN_R · Latent Space). The same author cataloged eighteen open clones before running the benchmark, and a Japanese thread notes the six earliest ones each guessed a different architecture: a 421M ModernBERT encoder with PPO, a diffusion reproduction, a Qwen3.5-9B LoRA, a three-class NLI classifier, a 40KB-embedding attention model, and a 0.5B LoRA that runs on a MacBook Pro. Nobody knows what is actually inside Jev, so the clone set is a spread of hypotheses. Extends the clone commoditization page.
- The counter-signal: Laya is not actually a Jev replacement (@kunchenguid · Laya). The strongest pushback of the slot, and it lands. Laya supports only 512 to 1K context, which rules out most real use cases, and measured directly its accuracy is close to a coin flip until you fine-tune it. This directly complicates yesterday's "the moat does not exist" reading in 2026-09-19-jev-open-clones-commoditization.md; an interesting research artifact is not a usable product.
- 938 of 1,284 tool calls were binary, and they cost 60% of the bill (@0xCarnagee). The best production number in the slot. Tagging every tool call in one agent session showed $38.61 of a $64.77 run went to yes-or-no questions (does this page link, is this a duplicate, safe to apply, keep or drop from context). Rerouted to a classifier the same session cost $26.16 at 38ms per decision versus 1.84s. Twenty-eight minutes of that run was a frontier model thinking about yes and no. Straight into the cost argument in llm-routing.md.
- Jev tooling ecosystem: eight repos in one thread (@zhouluobo · fast-jev-compaction). The most useful list of the day.
fast-jev-compactionis a Claude Code plugin that decides after every tool call what to keep, compress, or drop from context, which is context-window engineering done by a classifier instead of an LLM. Alsojev-browser,jev-search,pi-jevfor tool-call control, andhermes-jev-skillswiring the decision layer into model routing, memory, and compaction. Relevant to agent-harness-engineering.md. - LLM2Jev: turn any HuggingFace causal model into a decision engine, no retraining (cluster of 2: @_avichawla · @NFT_Chen · repo). The mechanism is next-token scoring over a constrained option set, prefill-only with no autoregressive decoding, served through Transformers or SGLang. This is the cheapest possible version of the idea: you already have the model, you just stop letting it generate.
- laya-mlx: 60 decisions per second on an M3 Max in under 1GB (@mizorewww · repo). An MLX port claiming 50x faster than hosted Jev with 7 to 14ms short decisions, demoed by playing Snake locally. Worth weighing against the accuracy caveat two bullets up: fast and wrong is still wrong.
- What Jev actually is, explained without the hype (@BenjDicken). The clearest plain-English explainer to circulate so far. Input is a text state plus a set of typed questions, output is a full JSON of answers produced in one pass rather than token by token. If you only read one Jev post from this slot, read this one.
- Jev as a dedicated decision layer in the agent loop (cluster of 3: @0xRicker · @yuhasbeentaken · @suraj_sharma14). Three independent posts converging on the same architecture: LLM reasons, classifier decides, code enforces thresholds and permissions. The ranked use-case list is the practical one, led by model routing, tool-call risk gates, intent triage, and agent loop stop/continue decisions. The 193x faster and 444x cheaper figures are self-reported and should be treated as marketing until someone reproduces them.
- A classifier is a judge, so Jev works for evals (@HamelHusain · blog). An LLM judge is a classifier, so a decision-only model can replace it, with the caveat to test against human labels and not overfit. The linked argument for binary pass/fail over 1-to-5 Likert scales is the substantive part and pairs naturally with a model that only emits typed choices.
- classifier.dev: zero-shot classification over plain HTTP, no API key (@DataChaz · site). Built on Jev with a smart tier that re-asks whatever the fast tier was unsure about, which is confidence-threshold routing in its simplest form. Published numbers on 400-item sets: 87.5% fast versus 90.0% smart on AG News. The framing that "a former OpenAI researcher spent two years and got open-sourced in three days" is overstated given the accuracy caveats above.
- Tuning an off-the-shelf decision model with GEPA instead of retraining (@richardartoul). Jev ships trained only on synthetic data, so adapting it to your criteria means prompt optimization plus a handful of labeled examples rather than a fine-tune. A cheap path to domain fit if it holds.
- Bespoke Nimble, an open model matched to Jev (@aisearchio · repo). Jev is API-only and closed, Nimble is the open version, self-reported at 90% against the original's 93%. It also ranks on JevBench, so this claim has an independent check available.
- Decision models for robot control, benchmarked in MuJoCo (@openroboto). Jev against GPT-6 Astra and GPT-4.1 mini on a pick-and-place task, each model choosing intent then X/Y/Z direction and gripper state. Low-latency typed decisions are a natural fit for a control loop, and this is the first attempt to measure it rather than assert it.
- Two Jev demos that are fun and prove nothing (cluster of 2: @nailthy62 · @MoonGotchi). A real-time virtual try-on at $0.0011 and 620ms per decision, and an autonomous onchain trading bot that has lost its builder $31,680 so far. The trading bot got 582K views and the highest engagement in the slot, which is exactly why engagement is the wrong ranking signal.
- Jev hype without content (cluster of 3: @AIGuide_ · @DataChaz · @jpthor). Reaction posts and memes with no claim attached. Skip.
- Opaque X articles, click through to read (cluster of 3: @Layton_Gott on Jev use cases in agent loops, @imryven on graph engineering after an agent lied about checking its own work, @_avichawla on building a local decision engine). Native X long-form, so the previews stop mid-sentence. The graph engineering one looks the most substantive of the three.
- Noam Brown on what actually solved Navier-Stokes (@AI_Whisper_X · Dwarkesh interview). The concrete numbers: 10,000 agents, 130 billion tokens, 88 hours, which Dwarkesh converts to roughly 4,000 years of one person thinking. Brown's framing is that multi-agent is just turning serial test-time compute into parallel compute, and he credits it under 10% of the result, with the model itself doing the work. Notably OpenAI gives agents one primitive, messaging each other, with no coordinator. Ultra Mode defaults to 4 agents for roughly 2x speed at 2x cost. Bears directly on multi-agent-systems.md and test-time-compute-allocation.md.
- GPT-6 Astra credited with the primary idea on a FrontierMath problem (@rynorhn). The researchers attribute the proof idea to the model and say they doubt they would have found it. Worth tracking as a claim rather than accepting, since "attribution" in these writeups has been elastic before.
- A CUDA guide worth the read (@gpusteve). Part of an AI performance engineering repo, pitched as putting you ahead of 90% of people on CUDA if you work through it. Feeds gpu-kernels.md.
- Serving 100 fine-tuned models on one GPU (@_avichawla). A repost on LoRA multi-tenancy, avoiding a separate endpoint per fine-tune. Standard material but squarely on the serving-cost axis.
- Suleyman is "really concerned" about Anthropic's model-welfare language (@rohanpaul_ai). Microsoft's AI CEO objects to Anthropic's constitution speculating that Claude may have preferences, feelings, or moral welfare, citing the retirement interview conducted with Opus 3. A genuine philosophical split between two frontier labs, not just posturing. Relates to responsible-ai.md.
- The safety-is-regulatory-capture argument, from three directions (cluster of 3: @VaibhavSisinty · @sovereignbrah · @GaryMarcus · Substack). Palantir's Karp argues on CNBC that the safety movement is a bid for nationalization, because labs absorbed enterprise proprietary data and only the US government can shield them from the lawsuits. A second post makes the commercial version: open weights deliver most of the capability at a fraction of the price, so incumbents want competitors banned. Marcus takes the personal-credibility angle at Amodei. Three different framings, one thesis, and none of them offer evidence beyond motive.
- Marcus on agent security, three years later (@GaryMarcus). Reprising his 2023 warning that giving agents unrestricted read/write internet access would be a security nightmare. The victory lap is tiresome; the underlying point tracks with this wiki's recent harness-permissions material.
- The enterprise bottleneck is context, not model intelligence (@rohanpaul_ai). Databricks CEO Ali Ghodsi argues most organizations are nowhere near the frontier of adoption, and what is missing is meetings, decisions, workflow history, and institutional knowledge in employees' heads. A useful corrective to the capability-race framing. Relates to agent-memory.md.
- Formal verification is being rediscovered badly (@RosuGrigore · blog). The argument: every hot field, blockchain then AI, claims correctness because "we use Lean," without a formal semantics of the language being verified. Garbage semantics in, garbage verification out. A sharp and underweighted objection to the verified-agent-output direction.
- LLMs Can Self-Improve, the 2022 paper (@hooshaaii · arXiv). Chain-of-thought plus majority voting across high-temperature sampling to filter confident rationales, then fine-tune on the self-generated data. Four years old and still the ancestor of most current self-improvement work, which is why it keeps resurfacing.
- Token pricing anxiety (@karanvasudeva). Pushes back on the claim that developers will be priced out of tomorrow's frontier models: today's smart models will be tomorrow's cheap ones, and software shipped fine before any of them. Commentary, but it is the right question about where the cost curve actually bites.
- The brain as four competing systems, not one (@IntuitMachine). Stanford work on the forebrain and hindbrain having separate embryonic origins, framed as an argument against a central command center. Interesting neuroscience, thin as an AI architecture analogy.
- Small tools and engagement bait (cluster of 4: @KanikaBK on an open-source Qwen-powered video clipper, @tomzaragoza on a hosted MCP server for Google Search Console, @IamKhanPhD on "15 Claude prompts" for PhD research, @manish_fp). The MCP server is mildly useful if you do SEO work. The rest is save-this-thread farming.
- Off-topic and promotional (cluster of 7: @GaryMarcus "AGI FTW @ FT", @SaveAbusedCows, @airindia, @BhuvanaGroup, @desi_dime, @bitget, @timesproindia). Ads, a charity drive, and a link-free one-liner. Skip.