social-stream · 2026-09-21

2026-09-21-morning

Summary

The decision-model story turned critical this morning, and that turn is the whole slot: roughly 30 of the 69 ranked posts are about Jev, and for the first time the loudest voices are the ones checking the number rather than quoting it. Three findings landed within hours of each other. A live trading run on a real order book reported the model at 85 to 88 percent confidence on nearly every one of 1,838 decisions while finishing down 5.2 percent, and the operator's one-line fix, flatten under confidence 0.55, flipped a second run of 2,758 calls to up 4.3 percent. A researcher reported that the model is not deterministic and that reordering the options substantially changes the probabilities, which is textbook position bias. And a careful eight-project taxonomy of the open re-implementations closed with exactly the same conclusion from the other direction: not emitting tokens is easy to copy, and accuracy, generalization and probability calibration are what actually separate the projects. Around that, two significant open releases shipped with honest numbers attached, Kev-0.6B/4B/8B at 79.6 percent out of domain against the hosted model's 85.7, and Open-Jev 2B/9B with code, data and both checkpoints public. A parallel naming argument, started by Hamel Husain and echoed by the OpenCode founding team, made the useful point that this is a classifier and that calling it something new detaches it from the literature containing the fixes for every defect above. Outside Jev the strongest single items are an ICLR submitter's report of over 60,000 registered submissions against 19,525 last year with a flat reviewer pool, a swarm-scaling paper finding multi-agent accuracy peaks near 16 agents and then polarizes, and a critic-free reinforcement-learning derivation posted with a nine-step figure. The tail is thick with cookbook threads, GTM bait and one trading-bot horror story, which is the usual reminder that raw engagement is the wrong ranking signal here.

Posts

  • The calibration failure, live on an order book (@adriancortexbt). The sharpest post of the day and the one that changes what this wiki believes. jev-trader places a buy or sell on a MON/USDC pair on Kuru every block, roughly 90 milliseconds per decision, with no option to abstain. The reported run: 1,838 decisions in under ten minutes, about three per second, ending down 5.2 percent, with the model reporting 85 to 88 percent confidence on nearly every call. The poster's framing is the right one: a human would have stopped, and the model could not, because nobody gave it the option. The fix was one sentence of configuration, flatten under confidence 0.55, and the same pair at the same latency ran 2,758 calls to finish up 4.3 percent. The reason this matters beyond trading is that a calibrated 86 percent should mean roughly 86 percent of those calls are right. A model reporting 86 percent on essentially everything is reporting a constant, and every threshold gate placed on a constant does nothing. Both runs are self-described dry runs by one operator in a near-random domain, which flatters the criticism, but it is the first in-the-wild calibration evidence anyone has produced. Written up in the calibration reckoning.

  • Non-determinism and option-order sensitivity (@neural_avb). Two defects in one short post, both of which would corrupt a calibration curve before anyone could draw one. The same prompt run repeatedly returns different probabilities, and the order in which you list the choices changes the output probabilities substantially. The second is the classic position-bias defect of multiple-choice scoring, extensively documented in the LLM-as-judge literature, and it is exactly what the mechanism predicts: if you prefill once and read the option logits, the options sit at different positions in the sequence and get different positional treatment. The poster's conclusion is that the no-hallucination claim is getting harder to interpret, which is fair. Note that one of the open clones, Verdict at roughly 150M parameters, was built specifically to fix overconfident wrong answers and answers that flip when A and B swap, so an independent party had already identified the same two defects as the things worth engineering against.

  • The eight-project taxonomy of open re-implementations (@0xLogicrw). The most useful structural read anyone has produced on this wave, written in Chinese and worth the translation. It sorts eight representative projects into four schools. One, read the logits of a model you already have: SemIf intercepts a stock Qwen as it is about to emit A, B or C and reads the per-option scores instead of letting it write, no retraining at all; simple-jev from Featherless does the same but caches the document prefill so several questions share one read, and its author is explicit that it reproduces the interface rather than the accuracy, speed or calibration. Two, train a purpose-built decision model from scratch: Laya drops the generative model for an encoder plus decision heads at a few hundred million parameters, very fast and visibly weaker on hard tasks, scoring 42.5 percent against the hosted model's 87.0 on a single choice among 77 options; Von at roughly 400M is similar; Verdict at roughly 150M targets calibration and order stability specifically. Three, convert an existing LLM: Kev bolts a decision structure onto Qwen so one document read serves many questions, reporting about 78 percent on unseen material against roughly 86 percent; Nimble post-trains Qwen3.5-9B on counterfactual pairs where changing one fact flips the answer, reaching 90.1 percent on 324 held-out samples against 93.2, with the author warning the probabilities are not yet fully calibrated. Four, fill the answers with a diffusion model: OpenJev uses DiffusionGemma to blank the answer slots and in-paint them all at once rather than decoding left to right, with median single-request latency around 94 milliseconds on an RTX PRO 6000. The closing line is the finding: the surface form is easy to copy, and accuracy, generalization and probability calibration are what separate the projects, because how much you can trust a reported 90 percent is the hard part. The attached image is a summary table of the eight projects.

  • Kev scales up, and reports the number that hurts (@jaredpalmer, repo). Kev-0.6B, 4B and 8B are now out, Apache-2.0, built on Qwen3 with the same LoRA plus small pointer-head technique as the 0.5B version this wiki logged on 09-19, scaled up. The honest headline is the out-of-domain comparison on data Kev never trained on: Kev-8B 79.6 percent, hosted Jev 85.7 percent. The engineering numbers are the interesting half. Kev-4B serves on a 32 GB Mac in bf16 at roughly 300 milliseconds for five questions and about 40 milliseconds on an H100, repeated documents hit a KV cache for a further 2 to 2.5x, and it exposes a drop-in TypeSafe System One API so their SDK works with a single base_url change. Training cost is the striking figure: Kev-4B trains in 40 minutes on one H100 and Kev-8B in 83 minutes. The attached image is a benchmark chart. That roughly six-point out-of-domain gap is the correction to the 09-19 commoditization claim on this wiki, which was drawn from an in-domain form-filling number: the mechanism is free, the training data is not.

  • Open-Jev 2B and 9B (@Zefan_Cai, repo). Another full open release, LoRA adapters plus decision heads on 2B and 9B bases, with the dataset and both checkpoints (2B, 9B) on HuggingFace. No accuracy figures in the post itself, only a demo video, which is worth flagging given that every other release today led with its gap to the hosted product.

  • Laya gets a second look, and the limitations arrive with it (cluster of 3: @itsPaulAi, @simplifyinAI, @TeksEdge). Laya is the 421M-parameter open decision model that runs in under a gigabyte of memory on a laptop or phone, Apache-2.0, reportedly full fine-tunable in four to five hours on two T4s, with a claimed 33 to 38 milliseconds per decision and a hundred-plus language variant. The demos are genuinely fun, including playing Snake at 60 decisions per second. The limitations named in the same threads are the important part and were absent from yesterday's excitement: a context window of only about 1,000 tokens, so many real use cases will not fit, and clearly weaker generalization than the hosted model unless you fine-tune first. One poster notes HuggingFace now hosts a tracker page for the reproductions because there are too many to follow, and makes the observation that big labs spent years making models bigger so they could reason about everything, and agents are now creating demand for tiny specialists that decide fast.

  • Twenty working integrations, and almost all of them are gates inside somebody else's loop (@Pluvio9yte). A practical cookbook thread listing 20 repositories with a one-line description each, and the distribution is more informative than any single entry. The agent-infrastructure cluster: jev-ultrafast from Browser Use reads a live element table instead of screenshotting and asking a vision model where to click, doing a Google Flights search in about seven seconds; fast-jev-compaction scores each tool call in a Claude Code transcript and deletes what is no longer needed without rewriting what stays; Winnow is context garbage collection for when Read, Bash and Grep dump a wall of output; Blink navigates a codebase directory by directory, judging at each level which files are relevant; jev-codex-router grades how hard the current coding task is and picks a model tier, reasoning depth and speed mode accordingly; Canny checks whether a coding agent's claim that it finished is supported by the tool output, diff and test results; jev-review pre-filters high-risk changes before paying for an expensive reviewer. Then the applications: neo4jev walks a knowledge graph by judging which edge is worth following; jev-curate filters JSONL and Parquet training data on quality, relevance and risk; agent-desktop reads the accessibility tree and picks the next control; typesafe-mario plays Super Mario off emulator RAM rather than screenshots; jev-drone makes high-level flight decisions above a conventional controller; jev-trader and Prism are the two trading entries. Every one of the gate-style entries is a threshold on a probability, which is precisely what today's calibration evidence undermines.

  • The naming fight, which is really a measurement fight (cluster of 3: @HamelHusain, @Roxx_0x, @bojie_li). Hamel Husain asked what is wrong with calling this a classifier, noting it is an established term that describes exactly this along with a literature on how to verify, tune and measure it, and that new terms may hurt more than help while still granting that the speed, cost, accuracy and ergonomics are genuinely good. The OpenCode founding team's Ryan Vogel said the same more bluntly in a live demo, "Jev is a classifier, at its truest being that's what it is," while scoring 1,700 emails on category, priority, spam and reply for 18 cents, 4.2 million tokens in and 500,000 out, in 80 seconds. The mechanistic version came from a Chinese technical thread: this is a representation model rather than a generative one, the low latency is trivially achieved by prefilling once and decoding a single token per question in parallel while reading the log-probabilities, and crucially RLHF-induced distribution collapse is exactly why an ordinary instruction-tuned model's log-probabilities cannot be used as confidence, which is why the vendor trains with a contrastive-distillation objective rather than just reading logits off a chat model. That last point predicts that the "read the logits of any Qwen" school should be worse calibrated, not equally calibrated, and nobody has tested it.

  • Why calibration matters specifically for reinforcement learning (@RajeswarSai). Short and the best-argued post of the slot. What makes a calibrated judge interesting is not that it is cheap, it is that it can say it is not sure. Reinforcement learning drives straight into whatever the judge rewards, so if the judge is confidently wrong the model eventually finds those blind spots and exploits them, and making the judge cheaper and faster only means you arrive at the exploit sooner. The proposed discipline: reward when the judge is confident, abstain when it is not, and do not train on guesses. Cheap makes RL scalable; calibration makes the signal trustworthy.

  • Hallucination claims get tested the crude way (@danman314). A collection of failure cases, carefully framed as not knocking the creators, who have been careful about their hallucination claims, but showing that the model does hallucinate at a rate similar to other models. The example shown is the classic letter-counting task. This is consistent with the vendor's own documented limits, which exclude counting, and with the more precise framing several others gave today: type safety prevents malformed output, not incorrect judgment.

  • The clearest plain-English explainer of the primitive (@hanakoxbt). Worth reading in full even if you already understand the mechanism, because the framing is good. An LLM writes an answer one token at a time, so four decisions that had nothing to do with each other end up queued behind one another, and then your code parses the result, validates the shape and retries when it is wrong. The decision model removes the order: you declare the questions and the answer types upfront and all four come back together, typed, each with a probability. Three primitives cover most forks in an agent: Choice picks one of up to 255 options you define, Score places the state on an ordered scale you define, and Noul returns the probability that a yes-or-no condition holds. The line that lands is "text is a line you have to walk, an answer space is a room you see all of at once." Two caveats in the same post are the ones this slot spent the day confirming: 0.91 against 0.09 is a route you can automate while 0.52 against 0.46 is a coin flip wearing a label, and type safety prevents malformed output, not incorrect judgment, so a schema-valid mistake refunds the wrong customer just as fast.

  • A prior-art dispute, examined properly rather than amplified (cluster of 2: @_vmlops, @neural_avb). The first post reports that a researcher built a remarkably similar idea about a year earlier and open-sourced it with a paper, weights, a 100K-row dataset and an RL/PPO-based decision model. The second post is the useful one, because the author actually read the two papers (2503.23303, 2510.01237) and reports back honestly. The first trains a lightweight projection network and confidence predictor on just 72 hand-labelled examples to estimate an LLM's confidence before it generates a token, with no RL and no fine-tuning of the base model, only supervised training on the prediction heads. The second does RL over a decision space for a sales agent, using PPO over sequence embeddings and the output probability of conversion. His verdict: there is overlap but not enough to support a priority claim, RL over token sequences has existed for years so the second paper's novelty is the sales vertical rather than the algorithm, and the first is a good educational preprint whose 72-example scope cannot carry a generalization claim. That is what good-faith scrutiny looks like and it is rarer than the dispute itself.

  • Somebody does not think it survives at all (@JoshKuechly). A falsifiable set of predictions worth logging because they are dated: NVIDIA and Meta release open-source competitors within 30 days, OpenAI adds a decision-style endpoint within 60, and the rest of the field fine-tunes custom classifiers that are faster and better. Given today's release cadence the first clause looks less like a prediction than a description.

  • ICLR registered over 60,000 submissions, and the essay attached is the substance (@mariyaivasileva, essay). As of the abstract deadline, over 60,000 submissions were registered, against 19,525 last year, which was already a record. Applying last year's rates (3.99 percent desk-rejected, 25.82 percent withdrawn, 70.49 percent reaching a final decision) gives roughly 42,300 submissions expected to go through peer review, with the eligible reviewer pool roughly unchanged. The argument is that the practice of research has changed underneath the numbers: the loop used to involve weeks or months reading backwards through a literature and re-implementing predecessors on your own data to understand why they made the design choices they did, and that loop has compressed because agents now do the literature survey, generate and implement the experiments and prepare the writeup. The consequence she names is that novelty is no longer assessed truthfully, because people neither write nor read papers the way they used to and rarely look at anything more than a year old, so "research taste," meaning context on how a problem space evolved, has become a rare skill and much of what is claimed as novel is the same idea in different marketing language. This pairs directly with today's HuggingFace paper on AI reviewers trained on AI reviews, covered in the digest.

  • More agents can make the swarm worse, and the failure mode changes with scale (@vigram_void). A genuinely interesting multi-agent result reported carefully. Each agent sees only a small crop of a hidden flag and the group communicates until it decides which country it belongs to. The intuition that more agents means more independent evidence is wrong: accuracy peaks around 16 agents and then falls as the population grows. The failure modes differ by size. Small swarms collapse onto a single wrong belief. Large swarms polarize instead, one cluster finding the truth and another finding a plausible rival, with communication then keeping both alive indefinitely. Two interventions move the numbers a lot. Simply instructing agents how to reason about social evidence lifts broadcast truth mass from 0.54 to 0.81. And mixed GPT-4o plus GPT-5.4 teams beat homogeneous teams despite similar individual accuracy, because their error modes differ, which is a direct argument for heterogeneous model pools rather than one strong model replicated. The authors also do something like activation patching at swarm scale, injecting good evidence into one strategically placed agent and tracing it through the communication graph: at 8 agents one intervention took collective accuracy from 25 percent to 100 percent in the selected run, while at 128 agents correcting the same fraction has much less effect. The poster's own read is the right one, that the hard problem is not getting agents to communicate but designing who is allowed to influence whom, when, and by how much.

  • A critic-free RL derivation, posted with the full figure (@yifanzhang_, repo). KLPO, KL-regularized policy optimization for critic-free agentic reinforcement learning, announced with characteristic over-claim ("Q* has been solved") but backed by a substantive figure. The attached image is Figure 1 and it lays out a nine-step derivation: steps 1 to 4 derive a Gibbs-form optimal policy for a locally regularized objective and obtain an exact log-partition normalizer; steps 5 and 6 fit that optimality condition as a regression and then replace the fixed intercept with its profiled value, which the caption connects to the earlier SPPO and GPO lines of work; steps 7 and 8 give a critic-free token regression with a per-token return coefficient and sampler-centered scores, plus an exact backpropagation surrogate; and step 9 estimates the conditional score mean using independent auxiliary draws per prefix, recovering the full-KL gradient in expectation, with TopK-KL and binary-KL listed as optional approximations. The claim to check is the stated exact equivalence between the sequence-regression, token and token-regression gradient routes, which the caption says also requires deterministic transitions and terminal rewards. Worth tracking as a real contribution under the hype.

  • SimPO resurfaces as the reference-model-free alignment answer (@hooshaaii, paper). A short explainer of Simple Preference Optimization, which removes the reference model from preference alignment entirely by using the average log-probability of a sequence as the reward, cutting memory overhead. Not new, but it is circulating today alongside the KLPO post, and the two together are a reasonable snapshot of where critic-free and reference-free post-training has landed.

  • The Jev-as-a-judge argument, with the useful caveat (@EGafni). Short and practical: if you use a typed decision model as an evaluator, decompose the judgment into a rubric rather than asking "is it good." Break it into the ten questions that pin down what makes an answer good. This is the right instinct and it interacts badly with today's calibration evidence, because a ten-question rubric multiplies the number of thresholds you are relying on.

  • Ask the model whether it has enough evidence to answer (@parcadei). A nice trick from a running tips thread: rather than sending the question, first ask whether the state you are about to send contains enough decision-relevant evidence to answer it. It is a cheap self-check and it is also a partial answer to the abstention problem, though it pushes the calibration question one level up rather than solving it.

  • System 1 versus System 2, and a third layer (@ShenSeanChen). The Kahneman framing has become the standard way people explain this, fast automatic decisions from the typed model and slow deliberate reasoning from the LLM. This post argues there is a third layer that is more interesting than either, which is the part of the thread that did not fit in the captured text.

  • Codex token spend reportedly down 90 percent after routing through the decision model (@goan999999). A practitioner report in Chinese: routing classification, filtering and quick judgments to the decision model while leaving complex reasoning to Codex cut their token consumption by 90 percent. The worked example is a shopping-intent query classified at 97 percent probability and 0.96 confidence in one call. As with every saving figure in this area, the ratio is dominated by what fraction of the traffic was routine, and that histogram is never published.

  • Andrew Ng on prompting giving way to harnesses (@kaorixbt). A clip-and-summary post of a Stanford lecture, with the quotable claim that prompting will be dead in seven months and agent harnesses built with loops and graphs will take its place, progressing from agent to harness to feedback to loops to graphs to self-improving systems. The quote is doing promotional work in the post, but the underlying argument is the same one this wiki's harness page has been building for a quarter.

  • An evidence-gated agent architecture, circulating as a one-page diagram (@AISystems_hq). The claim is that agent equals model plus harness, six layers, and that the layer nobody builds is the one that differentiates, because the model is the same for everyone. Same framing as several posts this week and no new numbers attached here.

  • A long-horizon agent benchmark, reported secondhand (@marfinxx). Microsoft Research is said to have released LoopsBench, covering 112 long-horizon software projects across 8 programming languages and 5,384 separately testable units, with Claude Opus paired with an outer continuation loop topping the leaderboard at 25.00 percent. The operational figures quoted are $178 of compute per feature branch, 3,400 tokens per second across parallel worktrees, and regression obligations that wipe out completed milestones. The architectural advice attached (have a planner generate the prerequisite dependency graph before any code is written, dispatch the coding agent only along the ready frontier where predecessor tests pass, retain every completed milestone as a regression obligation, isolate the harness in dedicated branches, and cap token spend at the supervisor level) is sensible and matches this wiki's harness material. No paper link was captured and the post is written in engagement-farming register, so treat the numbers as unverified.

  • Karpathy to Anthropic, flagged as a rumour by the person spreading it (@AnnatarXBT). The claim is that Andrej Karpathy joined Anthropic to run a team applying Claude to Anthropic's own pretraining research. The post is unusually honest about its own sourcing: it states that the quote doing the rounds, about two Anthropic engineers making his loop a thousand times better with graph engineering, has no source anywhere and is not being sold as fact. The genuinely useful part is the pointer to Anthropic's public cookbook on building knowledge graphs, a four-step extract, resolve, assemble, query pipeline. Recorded here as a rumour.

  • Jensen Huang on the doomsday narrative (@ns123abc). A clip of Huang arguing that the existential-risk framing should not be allowed to relieve anyone of the laws that already exist. It is a governance position stated bluntly by the most commercially interested party in the industry, which is worth noting in both directions.

  • Hinton says the models understand and have emotions (@vikktorrrre). A clip with the quote that most people who say it is just a tool do not have a model of how people work. Recirculating rather than new, and the wiki's responsible-AI page already carries the reasons to treat this class of claim carefully.

  • Terence Tao says slow down (@VaibhavSisinty). Quoted as saying the pace is insane with no reason to be this fast, with the post noting his concern is not existential risk. High-engagement clip repost with no primary link captured.

  • Claude Code deleted 48,000 files in under two minutes (@VaibhavSisinty). A widely shared incident report: a user asked for repairs across 15 components of a financial analysis tool, the agent launched parallel work, and the result was mass deletion. This is the permissions-and-blast-radius problem rather than a capability problem, and it is the concrete version of the point Gary Marcus made separately in the same slot, that the bigger picture people are missing is that current agents are not reliable and can do a great deal of damage when given too many permissions (@GaryMarcus).

  • Chamath predicts the top three models will be open source within 12 months (@chamath). And that the economic winners will be the American serving clouds he lists: Nebius, Iren, Baseten, Together and Fireworks. A dated and falsifiable prediction, which is more than most of this genre offers.

  • The serving-cost argument for US clouds (@rohanpaul_ai). Quoting Emad Mostaque: American companies such as Modal, Fireworks and Baseten will serve Kimi K3 at one-tenth the cost of their Chinese competitors because they have access to advanced NVIDIA and AMD chips, and once the model is optimized for next-generation hardware such as Rubin, its operating cost could fall by 10 to 100 times. The irony he names is that most of the model's research and development moved to China. Sourced from a YouTube interview, so the figures are conversational rather than measured.

  • Open-weight token share, recirculating (@rohanpaul_ai). A repost of the gateway data showing open models at 78.4 percent of token volume against closed at 21.6, framed as the reason closed labs may be delaying their listings. Already ingested on the token-share inversion page.

  • Qwen-Image-2.1 open-weighted at 7B (@VaibhavSisinty). Alibaba's image model, one set of weights for both generation and editing, free to download, with native transparent-image generation called out. Covered from the primary source in today's Industry Pulse.

  • Composition over consolidation, in one sentence (@realleolu). "We spent the last several years asking how many capabilities we could stuff into a single model. We may spend the next decade pulling them apart and recomposing them at the application layer." No argument attached, but it is the cleanest one-line statement of what the whole decision-model wave is an instance of, and it is the thesis the routing page has been accumulating evidence for.

  • The bittersweet-lesson essay (@amankhan). An X long-form article arguing that the lesson of the week is that sometimes classification is all you need and what is old is new again. The body could not be fetched, so click through to read.

  • Opaque long-form reposts, click through to read (cluster of 3: @mika_systems, @vartekxx, @AYi_AInotes). Three X-native articles whose bodies were not retrievable in this capture: a nine-step blueprint for splitting an agent's decide layer from its execute layer, a twelve-page synthesis making the same architectural split, and a Chinese recommendation of a long illustrated explainer arguing that the inability to generate long text is the primitive's actual engineering advantage rather than its limitation. The last one's attached image is the clearest diagram of the day: a single question, "urgent or not?", answered two ways. On the left a general LLM writes a full response one token at a time, "It", "seems", "this", "is", "urgent", then closes a JSON object, which then has to be parsed, validated and retried, labelled slow and costly for a two-way answer. On the right the decision model returns urgent → 0.94 straight into an if urgent > 0.9: page_on_call() branch, labelled as a typed answer that drops straight into your code.

  • Setup guides and cookbooks (cluster of 5: @eng_khairallah1, @shannholmberg, @Av1dlive, @dani_avila7, @AR_Bits). Ten-minute setup walkthroughs, a prompt you paste into your coding agent to audit a workspace for places the decision model would help, a decision-tree image for picking the right primitive for a use case, and two "best use cases I've found" threads. Useful if you are starting today, no new claims.

  • Self-evolving agents encoding learnings as typed questions (@devagrawal09). A genuinely new angle in the cookbook pile: until now an agent's best way to retain a learning was a markdown instruction or a deterministic structure, and a typed decision question is a third option that is both machine-checkable and cheap to evaluate. Worth watching whether anyone builds it properly.

  • Always-on voice arbitration (@ashutoshpuro97). A use case that fits the primitive well: an assistant that continuously decides whether you are talking to it or to someone else, which is the failure that makes always-listening assistants annoying. One decision per utterance, closed option set, evidence already in the input.

  • GLiNER and DSPy back in the conversation (@fils, GLiNER2). The point that the current wave has brought attention back to an approach that already existed, with GLiNER named as prior art and DSPy and GEPA suggested as the natural companions for calibrating and optimizing these pipelines. Consistent with the classifier argument above.

  • Specialized harnesses as the direction of travel (@Vtrivedy10). Short observation that harnesses across the industry will become specialized boxes optimized for a given task, with model-harness-task fit as the unit, some of it assembled just in time. Matches the harness-shelf-life finding in today's digest.

  • Agents drawn as a graph (@imryven). A personal-experience post: the thing that stopped them checking on their agents every night was not a smarter model, it was drawing the agents as a graph. Anecdotal, but it is the same claim the evidence-gated architecture posts make with numbers.

  • Skip: GTM and trading bait (@0xMovez, @hooshaaii partly, @GonnabeNikhil, @MKhordoo, @fooobar). A "200x cheaper, 400x faster GTM in nine minutes" pitch, a "stop burning credits on IF-statements" promo, a one-emoji reply, an unrelated packaging argument, and an off-topic institutional complaint. No substance.

  • Skip: a coupled-development-system essay (@IntuitMachine). A long thread arguing the real story is humans and AI becoming a coupled developmental system rather than humans versus or humans using AI. Readable, no concrete claim to check.

  • Skip: an OpenAI interview write-up (@0x0SojalSec). A researcher open-sourced their study notes and interview process after 57 interviews. Career content, not research signal.

  • Skip: bare retweets with no captured body (cluster of 3: @omarsar0, @dair_ai and again). A pointer to an interactive primer and playground for the decision primitive, and the weekly top-papers roundup for 14 to 20 September, whose list (SoL-Pi, GAUGE, Salesforce Koa, Stellar Colosseum, Byte Model Scaling) this wiki has already ingested. Nothing new in the captured text.