Skip to content

Jul 14, 2026

RAG is a database problem, not a model problem

Your retrieval pipeline fails long before your LLM does. What three enterprise deployments taught me about chunking, freshness, and who should own the corpus.

Abstract grayscale composition of nested rectangular blocks.

9 min

215 hearts32 comments

Jun 30, 2026

What evals actually buy you

An eval suite is a contract between your team and reality. How to write one that survives contact with the roadmap.

Abstract grayscale composition of horizontal bars of unequal length.

7 min

156 hearts18 comments

Jun 12, 2026

Agents need org charts

Multi-agent systems fail the way companies fail: unclear ownership, no escalation path. Borrow from management science, not sci-fi.

Abstract grayscale composition of concentric rectangular frames.

6 min

98 hearts9 comments

Paper summariesAll papers

Model notesAll notes

Jul 02, 2026

Claude vs GPT for enterprise RAG

Same corpus, same retriever, 400 real support queries. Where each model hallucinates, and why grounding instructions matter more than the model choice past a threshold.

Abstract grayscale composition of two overlapping rings.

12 min

187 hearts

Jun 20, 2026

Open weights vs API for regulated industries

Llama on your own metal against frontier APIs: the real TCO math, the compliance wins that are actually wins, and the capability tax you pay for control.

Abstract grayscale composition of a ring and a disc on a horizontal rule.

10 min

143 hearts

May 15, 2026

Small models for agent routing

Does your orchestrator need a frontier model? I routed 10k agent tasks through models a tenth the price. The answer is mostly no — with two sharp exceptions.

8 min

96 hearts

The WireAll news →

02

Jul 21

The EU AI Act's next compliance phase has a date

General-purpose model obligations start applying, and the guidance finally distinguishes a model provider from a deployer. If you fine-tune someone else's weights and ship it internally, read the deployer section — most enterprise teams are in it and assume they aren't.

Reuters

07

Jul 07

An argument that 'agentic' has stopped meaning anything

The piece is unfair in places and right in the main: a term that covers both a retry loop and a multi-hour autonomous session is not carrying information any more. Proposes a taxonomy by failure mode rather than capability, which is the more useful cut.

@karpathy


Solution architect · technology enthusiast · lifelong learner

Field notes on AI, architecture and strategy.

I'm Sunder. I help teams work out where AI fits, adopt it for real, and architect the whole path from ideation to production. Beyond the work, I'm happiest breaking dense white papers down until they read plainly, comparing how the newest models really hold up, and writing about what I find — with a bit of photography and the odd drawing when there's time.

Portrait of Sunder K

What lives here

01

The Wire

What actually happened in AI this week — short, dated, sourced, and linked to whoever reported it first.

02

The paper library

Foundational and current white papers, each with a plain-language summary, a glossary of the jargon, and my take on what it changes in practice.

03

Model notes

Everything I write about models: single-model reviews, hands-on notes, and the occasional head-to-head matchup.

04

The blog

Essays on AI — architecture, models, and whatever's got my attention that week. Readers can react and comment.

Working through an AI decision? Happy to talk it through.