Achieving Fit.
Distinctions — Essay 05 Derek Bronston
The human brain executes complex cognitive tasks on roughly 20 watts, the energy needed to run a small lightbulb. Large Language Models (LLMs), the type of AI model that drives ChatGPT and Claude, consume orders of magnitude more at production scale. This is evidenced by the growing demand for electricity and data centers driven by AI workloads. For several months I have been working on the following research question: how do we reduce the energy consumption of an LLM without reducing its accuracy? Armed with a lot of theory and drive, I employed Claude Code as an agentic research partner. Typically when I use a Claude agent to write code, I define a clear set of specifications for it to work with. I then architect and plan with the agent, defining the implementation expectations. In this case all I had was an objective.
My initial assumption was that I could innovate the solution with the agent given it was delegated enough theoretical direction, free rein, and a clear goal. However, as the results began to manifest, it became apparent that I needed a different strategy. The agent repeated the things it knew, producing variations on its existing patterns, but not breaking new ground. I kept asking for more, but I got more of the same.
It took me a while to understand what was wrong. The agent wasn’t failing. It was doing what it was good at. The real problem was that I was misallocating the work. The agent’s job was to do the things that were hard for me: hold context across long stretches, recall patterns that I would not have surfaced on my own, execute transformations at speeds I could not match, and read volumes of papers that would have taken me months to consume. An AI agent can quickly execute rigor, I can rapidly innovate on its summarization. Once I figured that out, the partnership began producing things neither system would have produced alone.
Consider RSA decryption. RSA is the cryptographic method that protects most secure communication on the internet: when you visit a website with a padlock in the browser, RSA or a close relative is doing the work. It encrypts a message into a ciphertext that looks like noise, and decrypts (transforms) the ciphertext back into the original message using a private key (which is numeric) and a specific mathematical formula.
When the ciphertext arrives, the private key is available, and the machine has the energy to compute. In theory it will produce the plaintext. However, the transformation requires reducing uncertainty across every variable in the equation before it can complete the work. RSA is deterministic, so the reduction of uncertainty is binary. Every variable is either fully resolved or the transformation fails. Energy and information alone do not ensure transformation. Uncertainty must drop to zero for each of the required distinctions.
Music functions in a similar way. When a musician improvises a line over a chord, they are reducing uncertainty across key, chord, voicing, time, and rhythmic placement. Choices rely on reducing uncertainty across each of these for the line to be produced. RSA and music are vastly different domains. RSA follows a prescribed algorithm; it is deterministic. Improvised music does not. It is non-deterministic, and does not follow a singular path to an outcome. Both share a property the physicist Stephen Wolfram named: computational irreducibility. Their outputs cannot be predicted without running them. The decryption has to be computed. The line has to be played. What distinguishes them is determinism. RSA produces one output for any input. Music produces one of many possible outputs, shaped in the moment. Success in both requires uncertainty to drop below a threshold before the transformation can complete.
In my essay ‘Understanding Is Not Enough’ I make the argument that a system requires the alignment of energy and information to achieve a goal (transformation). When either of those two is not aligned, the system fails. This essay points to a third failure. A system can have the correct information, its energy directed at the right outcome, and still fail. If a system cannot reduce uncertainty fast enough, or accurately enough, to execute it fails.
If we accept this premise, when a transformation fails and the information is correct and the energy aligned, the question is no longer “what went wrong?” The question becomes: why was the threshold of uncertainty not reduced to the level required for the transformation to succeed?
This is where cooperation enters. A single system has bounded capacity. It carries the distinctions it has been trained on, evolved into, or was designed for. It cannot resolve what it does not know. A cell cannot make distinctions about market conditions. A spreadsheet cannot make distinctions about pain.
Cooperation extends capacity. When two systems cooperate they can reduce uncertainty neither could resolve alone. Your gut bacteria break down fibers your cells have no enzymes for. The bacteria make distinctions about molecular structure that your digestive enzymes cannot. In return, your gut provides an environment the bacteria cannot construct for themselves. Together, they accomplish a transformation neither could complete alone. The distinctions do not overlap. They complement each other.
When a human cooperates with a Llarge Llanguage Mmodel it is structurally similar. LLMs predict the next token, the numeric equivalent to the next chunk of linguistic output. Humans predict the next idea. The model can reduce uncertainty across recall patterns and combinatorial possibilities the human cannot easily hold in their working memory. Whereas the human can reduce uncertainty about which question matters, which output is grounded in the actual situation, and which response moves the work forward towards its goal. Neither system does exactly what the other does. Each is reducing uncertainty the other cannot easily reach. Both can innovate and evaluate. The cooperation is optimized when each reduces uncertainty the other cannot easily resolve, regardless of which role they happen to be carrying in the moment.
When the cooperation does not work between an AI and human, it tends to fail in one of two ways. The human treats the model’s output as resolved and rubber-stamps it. The output degrades as the human is no longer contributing to the reduction of uncertainty the model could not. The other failure mode runs the opposite direction. The model produces something the human cannot verify, either because it sits outside the human’s ability to evaluate, or because the volume exceeds what the human can comfortably consume, degrading the cooperation into noise.
Both failures share a similar structure. The systems are no longer complementary because one has stopped doing the work only it can do.
When cooperation works mechanically, each system possesses distinctions about the other system’s capacity. The human has to know what the model is good at, and where it fails. The human also has to know what they are good at and where they fail. Without those meta-distinctions, the cooperation cannot route uncertainty to the system equipped to reduce it.
Effective cooperation is not produced by good intentions or shared goals. It is produced by fit between the cooperating systems’ distinction capacities. The fit can come from three places. Evolution produces it through selection: the bacteria that survived were the ones whose metabolism fit the environment your body produces, and the bodies that survived were the ones the bacteria made more efficient. Neither side carries meta-distinctions about the other. The fit is in the selection structure. Engineering produces it through design: when I started working with Claude Code, no selection process had pre-shaped either of us. I had to learn what the agent could resolve and what it could not, what it would generate fluently and what it would fail at. Then I could direct the work accordingly. The meta-distinctions had to come from me, and they had to be made explicit. Improvisation produces it in real time: musicians cooperating on a tune they have never played together generate the meta-distinctions as they play, listening for what the others are resolving and adjusting what they themselves resolve in response.
This claim has independent support in the technical literature. The same structural move appears in Clarity Theory, my framework for how systems reduce uncertainty to execute transformations, in Friston’s theory of biological cognition, and in transformer attention. Karl Friston’s free energy principle frames biological cognition as minimizing prediction error: a brain reduces uncertainty about its sensory input by updating its model of the world. The attention mechanism in transformer architectures reduces uncertainty about which parts of an input sequence matter for predicting the next token. The mechanism is not a metaphor borrowed across domains. It is what these systems are doing.
As my experiment has progressed, the cooperation has resulted in a shape that neither myself nor the agent would have produced alone. I made real progress on my initial goal of reducing energy consumption, and subsequently made an unintentional discovery that I will discuss in a future article. Regarding the process, I learned what the agent could resolve faster and farther than I could, and what it was less likely to resolve no matter how I prompted it. The agent had not learned anything about me. I had to make enough of myself explicit that the routing worked. I directed the work toward the distinctions only I could make. The agent ran the distinctions only it could make. The result was not the agent’s, and it was not mine. It was the output of the collaboration.
Information has to be reduced into actionable signal before a transformation can execute. A single system can only reduce uncertainty across the distinction sets it carries. Cooperation extends that capacity, but only when the cooperating systems express fit. Without fit, alignment and energy are not enough.

