# The gap in human-agent communication

Thoughts · Oct 2, 2026

Over the last few weeks, I've noticed a trend in assistant design: language models are being shaped to communicate in a way we can actually take in. A lot of hot new services like [Instinct](https://instinct.com/), [Poke](https://poke.com/), [Muse](https://muse.ai/), [Grok Bot](https://x.ai/bot) and [Dots](https://chatgpt.com/features/dots/) are trying to mimic the quick, natural rhythm of texting. One of [Opus 5.5](https://www.anthropic.com/claude-opus-5-5#communication)'s most widely praised improvements was how much more natural its responses felt. I also recently came across [Echo by Fulcrum](https://x.com/fulcrum_inc/status/2105416317834817871), a writing model that's great at matching human styles.

Most frontier models still aren't good at naturally communicating their thoughts to us. I find it an annoyance to read through the verbose, complicated outputs of coding agents. They surface a lot of stuff I don't want and miss what I would've wanted. Think about what you'd get back if you hired a capable human to do the job. It'd be three lines and maybe a question.

It's hard to tell if the issue is the model misunderstanding what I want or something ingrained in its vernacular. If it's misunderstanding, that's a much harder problem. A solution would almost need a personalized agent that keeps learning from you. If it's ingrained, we might be able to see traces of it in the pretraining data.

## High quality data isn't all you need

If models picked up this way of talking from what they read, it should show up in what they read. Frontier labs don't publish their data mixes, but NVIDIA [published the one for Nemotron 3 Ultra](https://arxiv.org/abs/2606.15007). Here's a rough estimate of that mix, sorted by writing format. It's built on a lot of assumptions and guesses, put together by Astra:

**Estimated pretraining formats.** Percentage of tokens, classified by each whole training item’s format.

| Writing format | Share of tokens | Teaches the model to |
| --- | --: | --- |
| Reference & explanatory documents | 22.6% | define every term, and to assume the reader knows nothing |
| Questions, answers & worked tasks | 29.5% | show every step, then restate the answer |
| Code files & formal source | 12.1% | be exact, because every token has to do something |
| Articles, blogs & stories | 16.8% | have a voice, and a point to make |
| Lists, records & other formats | 14.3% | break everything into bullets, tables and headings |
| Social posts, comments & reviews | 3.6% | keep it short, for readers who already saw the thread |
| Dialogue & speech transcripts | 1.1% | take turns with someone who already knows the context |

The first two together are 52.1%, written for a stranger who needs everything explained. Roughly half of the estimated mix.

*Astra classified a random sample of about 14,000 items from NVIDIA’s [public pretraining data](https://huggingface.co/collections/nvidia/nemotron-pre-training-datasets). Very rough, low confidence.*

About half of it is material written to be complete: documents and worked answers where someone (or a synthetic rewriter) spent real effort covering every detail for a reader they'd never met. That's great data. It just teaches models to write for a stranger, when what you want is a collaborator who already knows the context.

Chasing this purer form of communication isn't easy. The data that would teach it best (i.e. text messages between people) usually contains sensitive information and is understandably held to high data security standards. But hard problems haven't stopped this industry yet. Opus 5.5 was a great step, and the next few generations will hopefully close more of this gap than people expect.

[All writing](https://arjunsahlot.com/writing)
