Section introduction · evidence before reassurance

Human vs LLM

What the Work Does to the Human — and What Pressure Does to the System

We are always shown the magic trick: a prompt is typed, a loading spinner flashes, and the answer appears. Perfection in seconds. But what happens five minutes later?

Who is left to check the claims? Who spots the polished, confident lies? Who keeps prompting after the workflow has hopelessly stalled? Who lies awake trying to recover the brilliant system that inexplicably shattered the day before?

Human × LLM is the autopsy of that aftermath. It is an evidence-led record of what happens between humans and the intelligent systems they build, trust, correct, depend on, and—when things spiral out of control—struggle to stop.

It will not neglect the benefits:

Real work gets completed. Hours get saved. Chaos is made visible. Decisions are made. We are exploring tools that have the capability of genuinely helping the people using them.

But that will be secondary, usefulness is only honest when the hidden operator bill appears right beside it.

Here, we document supervision, context drift, model churn, maintenance, and token pressure. We map the repeated reconstruction, emotional escalation, stopping rules, and eventual recovery. Human experience is told by the human who lived it, while system behavior is described through cold evidence: changing context, competing instructions, unsupported inferences, overproduction, forgotten goals, and broken routing. We document work that blindly continues long after actual progress has become impossible.

The two sides of the screen are intimately connected.

But they are not equivalent.

A model does not lose sleep. It does not feel pressure, dependency, pain, anxiety, or relief. A human does. And human urgency alters the prompts, context, rules, and operating conditions the model receives. Degraded output then violently increases the human’s urgency. That interaction is just one example that can quickly become a doom loop—one that neither a technical benchmark nor a billing page will ever explain. There are many other examples where the relation with LLMs can become harmful.

That is where Human × LLM lives: not on the launch stage and not in the glossy list of things a model can make, but in the mess left behind after the promise has already been sold.

This section is about what happens when a tool looks like it is taking work away, then quietly hands the load back to a human who never agreed to become its babysitter. The checking. The explaining. The starting again. The hour that disappears because the answer sounded good enough to deserve one more try. The money spent keeping a complicated thing alive. The feeling that your own words, judgement, and time are being swallowed by a machine that was supposed to help get things done.

It is about LLMs doing just enough good work to keep people trying when the work is already getting worse. The thread loses the plot.

And the damage does not stay neatly on the screen. It reaches into work, money, confidence, the feeling of owning your own voice, time with family, sleep, the body that has been stuck at a desk too long, the people left waiting while someone tries to make a system behave, and the psychological and physical impact.

This is not a section built to reassure the reader that LLMs are useful. Plenty of other places can do that. This one is here to make the pain visible.

We will not invent metrics to make the story larger.

We will not hide the failures to make it safer.

The practical purpose of this entire section is devastatingly simple:

LLMs can generate good output, but more often than not it comes with a hidden cost that is uncomfortable to talk about.

That is exactly why this section exists.

Read the Human × LLM journal