The WireAI moves faster than it changes. Short notes on what landed this week, what it actually means, and a link to the source so you can disagree with me.
July 2026
Jul 23, 2026
ResearchPublication
The write-up covers the harness, the sampling temperature and the seeds — the three things that make a benchmark reproducible and the three things nobody publishes. Worth reading as a template for your own internal evals, not for the leaderboard position.
Jul 21, 2026
PolicyRegulator
General-purpose model obligations start applying, and the guidance finally distinguishes a model provider from a deployer. If you fine-tune someone else's weights and ship it internally, read the deployer section — most enterprise teams are in it and assume they aren't.
Jul 19, 2026
ProductVendor
At the new per-token rate, stuffing 200k tokens of context costs less than the retrieval infrastructure it replaces for a lot of low-volume internal tools. It does not change the calculus at scale, which is where most of the argument actually lives.
Jul 16, 2026
InfrastructurePublication
Two years, a team of nine, and the number is public. The interesting part is the split: roughly a third on the platform, two thirds on data access and governance work that existed before any model was involved.
Jul 13, 2026
ResearchResearch
The task set overlapped with a public repository that appeared in pretraining data. Scores dropped several points on the cleaned set. A useful reminder that an agent benchmark is a software supply chain, not a measurement.
Jul 10, 2026
FundingPublication
Money is moving from model training to model measurement. That tracks what I see in engagements: nobody is short of models, everybody is short of a defensible reason to pick one.
Jul 07, 2026
OpinionSocial
The piece is unfair in places and right in the main: a term that covers both a retry loop and a multi-hour autonomous session is not carrying information any more. Proposes a taxonomy by failure mode rather than capability, which is the more useful cut.
Jul 04, 2026
ProductVendor
The interesting engineering detail is the fallback: it routes to a hosted model when the local one is not confident, and the confidence signal is a separate small classifier rather than the model's own logprobs.
Every item links to its original source, which remains the work of its publisher. Headlines and summaries here are my own paraphrase and commentary, not reproductions.