The hardest part of a chat interface was the graphics

myynnin esteet

I built my first conversational interface in the late '80s with a friend, on a Commodore 64. Graphics on the 64 were just too much work, so a text-based "game" was the easiest thing we could make.

I didn't design another one until 2025. This time the graphics were the hardest part again, but for the opposite reason. Personality, memorability and the other things users still need had to be built around the edges of the experience, because the main interface was mostly words. Every word seemed to carry a lot of weight. The experience came from the interaction, the language and the tone of voice. There was very little else to work with.

It scared me. A conversational interface feels like the machine finally serves you instead of the other way around, but it needs a kind of clarity that can't lean on graphic design at all. It has to be done entirely in language.

Where did the settings page go?

For about two decades software made its complexity visible. Capability grew and the interface grew with it: dashboards, menus, filters, configuration panels. Users learned to find their way through the layers.

A conversational interface shrinks all of that into a text field. The system behind it is just as complex, often more so. Business rules, edge cases, permissions and failure states are all still there, and the system still has to work out what the user meant, what scope they're talking about and whether they're allowed to do it. The user just doesn't see the decision tree anymore.

The interface has become smaller. The complexity behind it hasn't. In a graphical interface the complexity is laid out in space. In a conversation it unfolds turn by turn, and the effort of using the system moves from finding things to interpreting them.

What a greyed-out button used to tell you

A visible filter tells you that the list can be filtered. A disabled button tells you there's a limit somewhere, and a modal tells you where the current task begins and ends. Users read capability off the screen without thinking about it.

A text field says none of that. Users have to work out what the system can do from how it behaves, or from documentation, and the cost adds up a little at a time. People hesitate before typing. They over-specify and hedge, send a small test query to see where the edges are, and rephrase to avoid being misunderstood. It's sensible adaptation, and all of it uses up attention they came to spend on something else.

"Not that. I meant the other format."

So the system has to guess. The guess is often roughly right and rarely exact on the first try, so giving a command turns into a negotiation: "Not that." "More concise." "Use the previous dataset." "I meant the other format."

Each round is small. Added up, it's coordination work that used to be built into the structure of the interface and now happens in the dialogue, done by the user.

For a product team this is awkward, because it's hard to see. Friction in a graphical interface shows up in clicks and drop-offs. In a conversation it looks like hesitation and rephrasing, and eventually like people using the product a bit less without ever saying why.

When the box accepts anything

There's a lot of talk about magic in conversational systems. The system understands, adapts, anticipates. The ones that work best operate inside a clearly limited domain. They set boundaries and shape expectations before the first message is typed.

Without those boundaries, an empty text field promises more than the system can deliver. When expectations run ahead of reliability, people slow down, test the system and second-guess what it gives them.

Every session starts from wherever the user happens to be, often distracted or already tired. In the Cognitive Budget Model that entry state is the budget you have to work with. If the interface adds uncertainty before it delivers anything useful, the budget is gone before the value arrives. A clean screen with one input field looks efficient, but if the user has to simulate possible outcomes in their head before typing, it's already costing them.

A graphical UI shows the map. A conversational one keeps the map inside the system, which moves the responsibility there too. The system now has to reveal its limits while the conversation is already happening. It also has to recover when it misunderstands something without making the user even less certain. Poor conversational design rarely looks broken. It feels slightly unreliable and slightly effortful, and people respond by using it less.

What to measure instead of how nice the chat feels

The useful question for a product-led team is whether users make decisions faster. Do they reach value sooner, commit to actions with more confidence, hesitate less over time?

If the interface looks simpler but leaves users less sure of what will happen, decision speed drops, even when satisfaction surveys stay neutral. People who are already overloaded reward predictability far more than a minimal look. Clarity here comes from a narrower domain and a system that behaves the same way twice.

The Cognitive Budget Model

The Cognitive Budget Model

Your users start every session mentally exhausted. The Cognitive Budget Model reveals how to design for reality, not ideal conditions.

Why is the web so afraid of feeling?

Why is the web so afraid of feeling?

When music matters to the person arriving, silence says something too.

Before they read a word, sound can tell them whether they are in the right place.