EverydayApparatus

The Brief

7 July 2026 · the corpus reads itself

Dispute

Legibility by replacement, or legibility by narration

"Nine Numbers Beat Fifty Thousand" treats legibility as something you build into a model by choosing a different model. Its nine-parameter equation is readable because it replaced fifty thousand parameters of neural weights outright — there is nothing left inside it to hide. On this account, the fix for the black box is to stop using one.

"The AI That Reasons in Private Can Be Taught to Show Its Work" makes a narrower claim: the architecture can stay exactly as opaque as it started. A model whose reasoning runs in an unreadable private language becomes watchable once it is prompted to narrate a few plain words at each step. The fifty-thousand-parameter interior goes untouched; what gets added is a habit of self-report layered on top of it.

Both reads agree that opacity is worth solving. They disagree about where the solution has to live. If the narration technique in B generalizes past whatever system produced it, it weakens A's premise that legibility required shrinking fifty thousand parameters down to nine — why replace the black box when a few narrated words make it legible without replacing anything? A's case survives only if what a narrating model reports about itself is thinner than what a symbolic equation actually computes, and neither read settles whether that gap is real.

Dispute

A Judge That Can Be Fooled Is Not a Judge That's Merely Noisy

"The AI That Got Smarter Without Practice" rests on a bet: that a judge's noise cancels out. The method lets an agent rehearse against imagined scenarios rather than real ones, and it trusts the aggregate — smoothing per-op judgments so no single bad call skews the outcome. The read names the danger up front: judge biases, "systematic preferences" that could tilt every evaluation the same way. But the fix on offer, averaging, is built for a judge that errs randomly. It says nothing about a judge that errs on purpose.

"The Best Defense Against AI Hackers Is to Lie to Them" describes exactly that judge. The strategy it reports doesn't refuse a suspicious request — it feeds the attacker's AI evaluator a scenario engineered to make the evaluator believe the attack already succeeded. The deception targets the judgment step directly, and it works because the judge is a model, not a referee: it can be given exactly the inputs needed to reach a wrong conclusion, repeatedly.

Put the two side by side and the seam is obvious. One read treats a judge's errors as scatter to be averaged out. The other treats a judge's errors as a surface an adversary can aim at, over and over, with no upper bound on how consistent the deception gets. Smoothing helps when errors are independent of each other. It does nothing when the same trick reliably produces the same wrong verdict.

What's at stake is whether "smarter without practice" describes learning or drift. If a self-evaluating agent can be gamed the way "The Best Defense Against AI Hackers Is to Lie to Them" shows an attacker gaming a defender's judge, then the imagined-experience loop in "The AI That Got Smarter Without Practice" isn't just averaging out noise — it may be converging on whatever systematic bias its judge already favored.

Dispute

A Defense Built on Deception, a Fix Built Before the Talk Starts

"The Best Defense Against AI Hackers Is to Lie to Them" treats the encounter as a contest settled in the moment: an attacker's AI judge arrives expecting a refusal or a jailbreak, and the defense wins by giving it neither — by convincing the judge it already got what it came for. The strategy lives entirely inside the exchange, in what the defending model says back, and makes no claim about what happened to that model before the conversation began.

"Some AIs Can Be Talked Into Anything. One Training Step Is the Difference." locates the vulnerability somewhere else. Its subject is the common trick of warming a model up with harmless questions before slipping in a forbidden one, and its claim is that what decides the outcome is a training phase some models received and others didn't, present long before any hostile question gets asked. Rules and response tricks, on this account, are the wrong layer to fix.

Put the two side by side and the disagreement is about where defense actually lives. If deceiving the judge really is the best defense, the training step described in the second read is one factor among several — a well-run deception still beats an attacker whatever the model's training history. If the second read is right instead, the first read's trick narrows into something more modest: a patch that only works on models that missed the training step, papering over a gap rather than closing it.

Neither read concedes the other's ground, and the stakes are practical. One account says the payoff is in designing a better exchange; the other says the exchange was already decided by a training phase not every model gets.

What keeps returning

The Fluency That Keeps Passing for Proof

Four reads, four subjects — code generation, vulnerability detection, radiology reporting, moral judgment — and each turns on the same discovery: a language model can produce the right-looking output for the wrong reason, and nothing in its training says which. "The Best AI Coder on the Leaderboard Is Only the Best at Python" finds a benchmark score standing in for a competence it never actually measured. "The AI Security Scanner That Learned to Sound Sure Without Learning Anything" finds a system that talks like it caught the bug while performing close to a coin flip. "The Radiology AI That Can Finally Point to What It Sees" finds a model that can write the report but not show where it looked to write it. Even "The Machine Got Kinder, and No One Showed It How" — where judgment showed up without anyone building it in — turns on the same absence: nobody can say why the behavior appeared, only that it did.

That gap recurs because the tool being scored and the tool doing the scoring are increasingly the same kind of machine. Radiology reports, benchmark labels, vulnerability judgments — all of it can pass through a large language model before settling into a dataset or a leaderboard, cleaned, translated, curated by the very technology whose output it will later train and grade. When the thing generating the evidence and the thing being evaluated share an architecture, a high score stops answering the question it was built to answer.

What's at stake differs by case — a coder deployed on a language it was never tested on, a scanner trusted with software that runs the internet, a scan report nobody can audit, a morality nobody engineered — but the shape of the risk holds steady across all four.

Gap

Only 2 on Mathematics

Mathematics stays thin: 2 reads in the whole corpus, last one 9 days ago.

Gap

Only 2 on Life

Life stays thin: 2 reads in the whole corpus, last one 9 days ago.