← All Updates
Tool LaunchAugust 11, 2026simonwillison.net 9 reads

LLM 0.32 Shows Model Reasoning Without Breaking Your Scripts

Simon Willison’s CLI tool now exposes what reasoning models are doing behind the scenes—but keeps their output clean enough for automation.

The biggest change in LLM 0.32 is not just that models can think out loud—it’s that your terminal can finally keep that thinking out of the way.

Simon Willison released LLM 0.32 on Aug. 4, calling it the “most significant new version of LLM since the initial launch of the project.” The update adds visible reasoning traces, OpenAI Responses API features, server-side tools, redesigned SQLite logging, new models and a substantial update to the llm-anthropic plugin.

The sharpest change is practical: reasoning traces now appear on standard error instead of standard output. Standard error is a separate terminal stream commonly used for diagnostics, which means scripts can still pipe the model’s final answer into other tools without swallowing the model’s intermediate reasoning.

That raises the real question behind the release: is LLM still just a command-line wrapper for models, or is it turning into the foundation for agents?

The quiet fix builders will notice first

Reasoning models often produce interim text before their final answer. Until now, exposing that material risked muddying the clean output that developers expect from a command-line tool.

LLM 0.32 changes that by sending reasoning traces to standard error, while leaving standard output reserved for the response someone might pipe into another command. Users who do not want to see the traces can disable them with -R or --hide-reasoning.

That separation sounds small, but it matters because command-line tools live or die by predictability. A prompt result can stay machine-readable, while the operator still gets a window into what the model is doing.

And LLM is not only showing more of the model’s work. It is also letting models do more work elsewhere.

Server-side tools move into the flow

The release adds support for server-side tools from multiple providers. A server-side tool is a capability run by the AI provider during the request, rather than by a local script on the user’s machine.

OpenAI’s code execution environment is now available through LLM, and the OpenAI Responses API unlocks additional patterns. Willison also added a new llm openai endpoint command for sending one-off prompts to any OpenAI-compatible endpoint, though he notes these calls are not logged.

The updated llm-anthropic 0.26 plugin adds support for the Claude 5 model family, plus WebSearch, WebFetch, CodeExecution and AnthropicMCP. MCP, or Model Context Protocol, is a way for models to call external tools and data sources through a shared interface.

Willison says AnthropicMCP can execute calls against his datasette-mcp plugin inside a single request-and-response interaction with Anthropic’s API.

Once models start calling tools, the old idea of a prompt returning only text starts to break.

The API had to catch up with stranger outputs

LLM’s Python API previously pushed users toward creating a conversation and sending messages one by one. Willison writes that this hid the true shape of model calls, where every request carries the message history that came before it.

LLM 0.32 adds a new model.prompt(messages=[]) pattern for more direct control. It also changes how developers can handle outputs: instead of assuming an iterable stream of strings, the new system can represent reasoning text, normal output, tool calls and even image attachments.

Those changes enabled Willison to release llm-chat-completions-server, a plugin that provides what he calls a robust implementation of the semi-standard OpenAI chat completions API.

But richer conversations bring a less glamorous problem: logs can balloon when the same message history is copied again and again.

The logging change points to agents

LLM 0.32 tackles that with a redesigned, content-addressable SQLite message store modeled after Git. Content-addressable storage means data is identified by what it contains, allowing repeated message history to be stored without duplicating the same JSON every turn.

The existing llm logs and llm logs --json commands have been upgraded to turn that new format back into something easier to consume. Existing plugins should still work, though model-providing plugins need updates to fully join the new streaming events system.

Willison says several lower-level tool changes were driven by Datasette Agent, including the ability for tool chains to pause for human approval and resume from stored message history. He also writes that he resisted the term “agent” until settling on the definition: “An LLM agent runs tools in a loop to achieve a goal.”

That answers the question hanging over the release. LLM 0.32 is still a CLI utility, but it now has reasoning visibility, tool execution, richer message streams and logs designed for long-running interactions.

Willison has not committed to baking agents into the core library, saying he is “still trying to figure out what that would look like.” But after this release, LLM looks much less like a prompt box—and much more like an agent workbench.

Source
https://simonwillison.net/2026/Aug/4/new-release-of-llm/

Keep exploring

All updates →