Winking at a Touchscreen: The Interface Problem in Agentic AI
There’s a specific feeling that has inserted itself into my work routine, like too many mg/caffeine. I’ve spent years learning how to do something — really learning it, being shit, being not so shit, feeling shit but being OK, then “it depends” confident — and then one day you’re told: don’t worry, we’ve got it from here. You can watch if you like.
It’s not quite retirement. It’s more like being benched for a better player. And if (like me) you’re in product, design, or any role that used to involve making things with your hands and your judgment, Agentic AI is starting to feel like a 6’4 12 year old who can already dunk.
But, maybe, just maybe, that kid can be your best friend…
The Blackbox Problem
The first cut is the loss of visibility.
An Agentic workflow barely shows you it is working. It disappears into a loop of planning, tool calls, and self-correction, then surfaces with an output. Sure I can skim-read. I even open up analytics one in a while. But ultimately the kid is too fast.
So. You really only have two options: accept it, or start over. There’s no deep co-authorship here. There’s no equivalent of looking over someone’s shoulder and saying “wait, not quite”, or taking the pen from them to scribble it a different way. There is no bounce, no combined flow, the banter is it printing:
"Moseying… (8m 19s · ↓ 35.7k tokens)"
What.A.Hoot.
This isn’t a niche complaint. It’s the difference between a collaborator and a vendor. When you hand work to a vendor, you specify requirements upfront, wait, and review. Sometimes that’s fine. Often it isn’t, because you didn’t know exactly what you wanted until you saw the first attempt. Iteration is how thinking sharpens. Removing it doesn’t remove the need for it — it just relocates the error to later in the process, when it’s more expensive to fix. It is a lot less fun and, dare I say: it lacks any real agency.
The Volume Problem
Step under a waterfall and look up, open your mouth and try and swallow all the water that falls. That is your consumption limit. What you don’t drink is not yours.
AI systems are now generating approximately 310 trillion tokens per day globally — combining Google’s infrastructure (which alone crossed 107 trillion tokens daily as of May 2026, up 330x from April 2024), Chinese providers including ByteDance’s Volcano Engine, OpenAI’s API, and others.123
A token is roughly three-quarters of a word. So that’s around 232 trillion words generated by AI, every single day.
Sam Altman noted in early 2024 that all humans on Earth generate roughly 100 trillion words per day — spoken, written, everything.4 AI has already lapped us. It is generating more than twice the total word output of our entire species, daily, and accelerating.
Meanwhile, the average adult reads at 238 words per minute — the real figure from a 2019 meta-analysis of 190 studies — not the 300 WPM myth that circulates online.5 Across the world’s roughly 7 billion literate people, even if everyone spent 45 minutes a day reading AI output, humanity’s total reading capacity amounts to roughly 100 trillion tokens. We are generating roughly four times more than we can physically consume — and the gap compounds ~25% a month.6
The Interface Hasn’t Caught Up
An infinite number of monkeys, with an infinite number of typewriters are sending us love letters. Or hate mail. It is hard to tell when there is so much. My argument is not about stopping the monkeys though. I love their little chimpy notes. My argument is for a better postal system. I want the SMTP of infinite monkey letters. But we’re still in the “lick the stamp and post it” phase. The question isn’t whether the current chat-and-output paradigm is optimal. It obviously isn’t. The question is whether the switching costs are low enough for something better to take hold.
To tell it another way. Interacting with most AI systems feels like winking at a touchscreen. You’re making a gesture that can work, in a medium that doesn’t quite support it. You describe what you want, the system produces something, you describe the delta, it produces something else. The feedback loop is verbal when it should sometimes be gestural, visual, or spatial. The output is a wall of text when it should sometimes be a diagram, a graph, a button that says yes, exactly that.
Thariq Shihipar, engineering lead on Anthropic’s Claude Code team, made a version of this argument in May 2026 when he wrote that HTML had become Anthropic’s internal default for agent outputs, replacing markdown. The reasoning wasn’t aesthetic — it was functional. Richer visual structure reduces the cognitive work required to understand what an agent has produced and decide what to do next.7 He called it “the unreasonable effectiveness of HTML.” The core insight was older than his post: a picture tells a thousand words, and in agentic workflows, fewer words between human and machine is worth more than more tokens.
What I Actually Want
I’m not asking for AI to do less, it will do more. Pandora left her box on the bus and we all get to look inside. But if this kid is going to dunk on me, or these monkeys are going to write to me I want to be the player coach, I want spam filters, I want something better.
In no particular order, here is my list of (somewhat) unreasonable demands of AI interfaces going forward:
Immersive collaboration, not shuttle diplomacy. I want to work inside the same artefact the agent is working on. Not review its output and annotate it in a separate pass, but be present in the process — watching a structure emerge, redirecting in real time, contributing directly. The distinction sounds subtle. It isn’t. One is creative partnership; the other is procurement. In my head it is something like an interactive whiteboard with a graph.
Adaptive interfaces. Not adaptive in the brittle way of current form wizards — “you must answer these three questions before proceeding.” Adaptive in the fluid way of a good colleague who hands you a colour picker when you say the shade isn’t right, rather than asking you to describe the shade you want in words.
Multimodal interaction. The ability to say not quite by waving a hand, by drawing a rough shape, by speaking a tone as much as a specification. Language is rich but it is not the only register in which humans think. Agents that can only receive text prompts are losing most of the signal. I want to be able to shrug and say “Yeah” and it know that I am saying yes out of defeat, not victory.
Contextual interruptions. Agentic workflows run longer. They need to be able to surface decisions at the right moment — not pinging constantly, not disappearing for hours. The model should be: judge the risk, judge the urgency, judge my current focus, then decide whether to interrupt. That requires something more sophisticated than “ping me for everything” or “do your best”. I want this to be contextual to MY work as well. Look at my calendar, check if I’m busy, then send it. If it is that urgent, call me, or stop and wait.
Visual check-ins over walls of text. A flow diagram of what the agent has done and is about to do. A graph of the decision tree. A DAG of dependencies. These are not decorative. They are the difference between oversight and rubber-stamping. I do not want this to just be usage. I want it to be planning. Interactive. Easy to modify, copy, paste… generally input into.
The Point
There is a version of agentic AI that feels like early retirement — all the work done for you, nothing left to contribute, a slow atrophying of the skills you spent years building. And there is a version that feels like having an exceptionally capable colleague who knows when to ask, when to proceed, and when to show you something visual rather than write you a paragraph.
The difference is not the model. It’s the interface, the interaction model, and whether the system is designed to keep the human genuinely in the loop — not as a checkbox, but as a participant.
Some of this is emerging. Google has been experimenting with generative UI elements that adapt based on task context. Claude Artifacts has generated tens of millions of interactive HTML outputs.8 Every vertical SaaS is wrestling with these ideas. The platforms are beginning to understand that text output is a special case, not a default.
The talented young player is learning to pass, the monkeys are getting crayons, and maybe, just maybe we all get better jobs.
References
-
Google processes ~3.2 quadrillion tokens/month as of May 2026 — roughly a 7× jump on the prior year — per Sundar Pichai at Google I/O 2026 (up from ~9.7T/month in Apr 2024 and 480T in May 2025). Source: CryptoBriefing. ↩
-
China’s providers: total daily token consumption reported at ~180 trillion/day (Feb 2026), with ByteDance’s Volcano Engine alone at ~63 trillion/day (Jan 2026). Source: Robonomics (Substack). ↩
-
a16z & OpenRouter, “State of AI” (Dec 2025) — OpenAI’s API averaged ~8.6 trillion tokens/day in Oct 2025; OpenRouter alone passed 1 trillion tokens/day. Source: a16z. ↩
-
Sam Altman (Feb 2024): “openai now generates about 100 billion words per day. all people on earth generate about 100 trillion words per day.” Via Nathan Lambert, Interconnects. Source: Interconnects. ↩
-
Brysbaert, M. (2019). “How many words do we read per minute? A review and meta-analysis of reading rate,” Journal of Memory and Language 109 — meta-analysis of 190 studies (~18,000 participants) putting average silent reading at 238 wpm, not the oft-cited 300. Source: ScienceDirect. ↩
-
OfTech Explorations, “Interface Beats Model — the throughput-vs-reading wall” (2026-07-20) — recomputes the surplus from this essay’s own cited figures: AI ≈ 413T tokens/day (310T words ÷ 0.75) vs human reading capacity ≈ 100T tokens/day = ~4.1x, compounding ~25%/month. The “3x” in the body is conservative; the cross-over is already passed. Source: /blog/explorations-interface-beats-model/. ↩
-
Shihipar, T., “Using Claude Code: The Unreasonable Effectiveness of HTML” (8 May 2026) — the Anthropic Claude Code engineer argues richer visual structure (HTML over Markdown) cuts the cognitive work of parsing agent output. Via Simon Willison. Source: simonwillison.net. ↩
-
Anthropic, on making Artifacts generally available (27 Aug 2024), reported that “tens of millions of Artifacts” had been created since the June 2024 preview. Source: VentureBeat. ↩