I worked on speech recognition for many years, and it's an almost solved problem. That is an exciting thing to be able to say. It is also a strange one, because it feels like the field is finished and the problem I spent years on is gone.
Then at Interspeech 2025 I ran into a slide from Roger Moore. It was a pyramid with eight layers. Spoken dialogue, the part I had been working on, was the base. Above it were seven more: knowing someone is there, knowing they are talking to you, timing, doing things together, and at the top, having a reason to speak at all. Each layer builds on the one below.
The problem was never recognition. It was understanding, and that one is always there. Recognition was the first step up a pyramid, and the rest of it is the road to a more intelligent communicative machine.
1. The pyramid
Read it from the bottom. It goes from being able to talk to having a reason to.
A machine with no goals has no reason to speak. A machine that isn't present can't engage. A machine that can't coordinate in time can't take a turn. A machine that can't act with you has nothing to talk about. The words are necessary, and they are the smallest part.
Moore has a second idea that explains what the base feels like on its own. He calls it the habitability gap. As a system gets more flexible, usability drops, because people can no longer tell what it can do. He reads it as the uncanny valley, and says you only avoid it when "the visual, vocal, behavioural and cognitive affordances of an artefact" match one another. A perfect voice on a machine that can't engage is a mismatch, and people feel it before they can say why.
2. The ML behind the pyramid
Here's the part that made the pyramid feel like home. Every layer is an ML problem I already knew, applied to a bigger unit. The lineage is token classification, sequence tagging, sequence to sequence, translation, language model, masked language model, causal transformer. Token-level tasks label tokens. Conversation-level tasks label utterances, turns and threads. The models changed. The formulations didn't.
The map first, then each method on its own with the one equation that defines it. Throughout, \(x\) is the input sequence, \(y\) the output, \(h\) a hidden vector the model computes, and \(\theta\) the parameters.
2.1 Token classification: Label every token
The simplest formulation. Each position gets its own label, and the labels don't talk to each other. Encode the sequence, put a softmax over labels on every position:
\[p(y_i \mid x) = \mathrm{softmax}(W h_i + b), \qquad \mathcal{L} = -\sum_{i} \log p(y_i \mid x)\]
Read it left to right: a vector for each token, a distribution over labels from that vector, and a loss that makes the right label likely. That's named-entity recognition. At the utterance level the same head answers "does a new stage start here?", with \(y_j \in \{\text{boundary}, \text{no boundary}\}\).
Example. x = [I, love, New, York], y = [O, O, B-LOC, I-LOC]. Four tokens, four labels: O for outside any entity, B-LOC for the start of a place name, I-LOC for its continuation. The model scores York as I-LOC from \(h_{\text{York}}\) alone. It never checks what it gave New.
Token classification is a softmax on every position, trained independently.
2.2 Sequence tagging: Labels that depend on each other
Labels are rarely independent. A question is usually followed by an answer. A conversation that has reached its closing doesn't jump back to the opening. Sequence tagging adds a term for the transition between neighbouring labels and scores the whole label sequence at once. This is the conditional random field:
\[p(y \mid x) = \frac{1}{Z(x)} \exp\Big( \sum_{i} \psi(y_i, x) + \sum_{i} \phi(y_{i-1}, y_i) \Big)\]
\(\psi\) says how well label \(y_i\) fits position \(i\). \(\phi\) says how plausible the step from the previous label is. \(Z(x)\) normalises over every possible label sequence. Decoding is Viterbi, which finds the best path in \(O(n \cdot |Y|^2)\) instead of trying them all. The hidden Markov model is the generative cousin, \(p(x, y) = \prod_t p(y_t \mid y_{t-1})\, p(x_t \mid y_t)\), and it's what stage detection ran on before transformers: a handful of stages, a transition matrix, an emission model over utterance features.
Example. Same x. The path [O, O, O, I-LOC] is impossible, because an entity can't continue without starting, so \(\phi(\text{O}, \text{I-LOC}) = -\infty\) and Viterbi never returns it. At the utterance level: x = ["Where is the exit?", "Down the hall.", "Thanks."], y = [question, answer, closing]. \(\phi(\text{question}, \text{answer})\) is high and \(\phi(\text{closing}, \text{question})\) is low, so the tagger prefers the path a conversation would actually take.
Sequence tagging scores the whole label path, so a stage can't jump where the grammar forbids.
2.3 Sequence to sequence: One sequence becomes another
When the output is a sequence with its own length, the model generates it one element at a time, each conditioned on the input and on everything it has generated so far:
\[p(y \mid x) = \prod_{t=1}^{T} p(y_t \mid y_{<t}, x)\]
Translation was the first use. The same factorisation turns a conversation into a state diff, a thread into a summary, a state into a reply. The encoder reads \(x\) into vectors and the decoder produces \(y_t\) by attending back to them.
Example. x = I love New York, y = J'aime New York. The decoder emits J' given \(x\), then aime given \(x\) and J', then New, then York, then a stop token. Nothing says four tokens in means four tokens out. Later in the essay the same model takes x = previous state + new utterance and emits y = {"order_id": "A1234"}.
Seq2seq generates the output token by token, conditioned on the input and on itself.
2.4 Language models: Predict the next token
Drop the input sequence and the seq2seq equation becomes a language model:
\[p(x) = \prod_{t=1}^{T} p(x_t \mid x_{<t})\]
That's the causal, left-to-right model, and it's the one that generates. The masked language model trains the other way round. Hide some tokens and predict them from both sides:
\[\mathcal{L}_{\text{MLM}} = -\sum_{i \in M} \log p(x_i \mid x_{\setminus M})\]
It can't generate, but its hidden vectors \(h_i\) have seen the whole sequence, which makes them the better input to a classifier or a tagger. In conversation work the encoders label things and the causal models say things.
Example. Causal: x = I love New and the model puts most of its probability on York. Masked: x = I [MASK] New York and the model predicts love, using I on the left and New York on the right. The vector for [MASK] saw the whole sentence, which is why encoder vectors make better features for labelling.
Causal models generate. Masked models represent. Both are next-token prediction with a different mask.
2.5 The causal transformer: One model for every formulation
All of the above now run on one architecture. Attention lets every position read every earlier one:
\[\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^\top}{\sqrt{d}} + M\right) V\]
\(Q\), \(K\) and \(V\) are linear projections of the hidden vectors. \(M\) is the causal mask, \(-\infty\) above the diagonal, so a token can't see its future. The same hidden state \(h_t\) feeds whatever head you put on it: a softmax for classification, a CRF for tagging, the next-token distribution for generation. And a prompt turns any task into generation, \(p(y \mid \text{instruction}, x)\). That's why one model now does the whole map table.
Example. For x = I love New York, the query from York attends to the keys of I, love, New and itself. For love, the mask hides New and York. The weight from York onto New comes out high, so \(h_{\text{York}}\) carries "second half of a place name." Put a label head on it and you get I-LOC. Put the next-token head on it and you get a guess at what follows. Put Tag the places in: in front and the same model writes the labels out as text.
One hidden state, many heads. The head, or the prompt, picks the formulation.
2.6 Embeddings and online clustering: Who is talking to whom
Several people in one channel, and the first job is to split the stream into conversations. A message gets an embedding that mixes its text with who sent it and when:
\[e_m = f_\theta(\text{text}_m, \text{speaker}_m, t_m)\]
Each new message is scored against the threads still open in a window. Cosine similarity to the thread's running embedding is the baseline. A learned scorer that also sees the time gap and whether this speaker is already in the thread does better:
\[s(m, T) = \sigma\big( w^\top [\, e_m ;\ e_T ;\ \Delta t ;\ \mathbb{1}[\text{speaker}_m \in T] \,] \big)\]
The assignment rule is the online part:
\[T^* = \arg\max_T s(m, T), \qquad \text{join } T^* \text{ if } s(m, T^*) > \tau, \text{ else open a new thread}\]
and the thread embedding updates as a running mean, \(e_T \leftarrow \frac{n e_T + e_m}{n + 1}\). Evaluate it as clustering against hand-labelled threads, with variation of information or exact-match thread counts. In speech the same problem is diarization plus overlap.
Example. A team channel:
m1 Ann, 10:00: "Anyone up for lunch?"
m2 Bo, 10:00: "The build is broken again"
m3 Cy, 10:01: "Sure, noodles?"
m4 Ann, 10:01: "Which test?"
x is each message with its speaker and time. y is a thread id. m3 scores high against m1 because it answers an open question, so T1 = {m1, m3}. m4 is the interesting one. By speaker it belongs with m1, since Ann started that thread. By content it answers m2, so it joins T2 = {m2, m4}. Text alone gets it wrong. Speaker alone gets it wrong. Reply structure wins.
Disentanglement is argmax-and-threshold over a learned similarity, one message at a time.
2.7 Hierarchical sequence model: Stage per utterance
Stage is a label on an utterance, and an utterance is itself a sequence, so the model has two levels. Encode each utterance \(j\) into one vector, then run a sequence model over the utterance vectors:
\[u_j = \mathrm{Enc}(x_{j,1}, \dots, x_{j,n_j}), \qquad h_j = \mathrm{Seq}(u_1, \dots, u_j), \qquad p(z_j \mid h_j) = \mathrm{softmax}(W h_j)\]
\(z_j\) is the stage. Put the transition term from 2.2 over the \(z_j\) and it's a CRF at the utterance level. That's where the stage grammar lives: \(\phi(z_{j-1}, z_j) = -\infty\) for transitions the domain forbids. Per-person stage is the same model run on one speaker's utterances, and the gap between a speaker's \(z_j\) and the thread's is a feature in its own right.
Example. A support chat, x = five utterances in order:
- "Hi, my order hasn't arrived"
- "Order number?"
- "A1234"
- "I've reshipped it"
- "Thanks, bye"
y = [problem, clarify, clarify, resolve, close]. Each \(u_j\) is the encoded utterance, each \(h_j\) has read the ones before it, and \(\phi(\text{close}, \text{problem}) = -\infty\) stops the tagger reopening a closed conversation. Run it per speaker and the customer is still at problem while the agent is at clarify, which is normal, or at resolve, which means the agent moved on too early.
Stage detection is sequence tagging where the token is an utterance.
2.8 State tracking: Update a belief
This is the step I wrote about in Robotics 101 under state estimation, with utterances in place of sensor readings. Utterances are noisy observations. Keep a belief \(b_t(s)\) over states, predict how it drifts, correct it with what was just said:
\[b_t(s) \propto p(u_t \mid s) \sum_{s'} p(s \mid s')\, b_{t-1}(s')\]
\(p(s \mid s')\) is the transition model, the stage grammar again. \(p(u_t \mid s)\) is how likely this utterance is in that state. For a state with many fields the neural version is a recurrence,
\[s_t = g_\theta(s_{t-1}, u_t)\]
and the current version generates only what changed, \(p(\Delta s_t \mid s_{t-1}, u_t)\): the seq2seq model from 2.3, reading the previous state and the utterance and emitting the fields to update. Incremental, because it reads one utterance. Probabilistic, because it keeps \(b_t\) rather than a label. Typed, because \(s\) has fields that code can branch on.
Example. Same chat. After utterance 1 the state is {issue: missing order, order_id: ?, stage: problem} with a belief over stage of {problem: 0.8, clarify: 0.2}. Utterance 2, "Order number?", moves the belief to {clarify: 0.9} and leaves the fields alone. Utterance 3, "A1234", changes one field, so the diff model emits {"order_id": "A1234"} and nothing else. The tracker never re-reads utterances 1 and 2. What it needed was already in \(s_{t-1}\).
State tracking is a Bayes filter where the sensor is the last thing somebody said.
2.9 Intent: Classify the move, infer the goal
Surface intent is classification, \(p(a \mid u)\), with \(a\) from a fixed set of acts. The goal is never observed. It has to be inferred from the whole sequence of moves:
\[p(g \mid u_{1:t}) \propto p(g) \prod_{i=1}^{t} p(u_i \mid g, s_{i-1})\]
Which goal makes the moves so far most likely? The inverse-RL version says the same thing with a reward. Find \(r\) such that the observed moves are near-optimal under the policy it induces:
\[\pi^*(a \mid s;\, r) \propto \exp\big(Q_r(s, a)\big), \qquad r^* = \arg\max_r \sum_i \log \pi^*(a_i \mid s_i;\, r)\]
Social intent is the same machinery with a much wider hypothesis space, so the posterior stays flat. That is the formal way of saying it's always a guess.
Example. u = "Is there any way to get this faster?" Surface intent: request. That's the label. The goal needs the sequence: the customer mentioned Friday twice, asked about express shipping, never mentioned a refund. The posterior comes out around {needs it by Friday: 0.7, wants a refund: 0.1, other: 0.2}. Social intent: "I guess that's fine" is an acceptance with a hedge, and the best estimate is something like {accepts: 0.6, unhappy but conceding: 0.4}. It should stay that uncertain.
Surface intent is a label. Goal is a posterior over what would explain the moves.
2.10 Policy: Choose the next move
Given the state, the machine picks an action from a small set, \(a \in \{\text{speak}, \text{wait}, \text{acknowledge}, \text{clarify}, \text{yield}\}\). First by imitation, from human conversations:
\[\max_\theta \sum_{(s, a^*)} \log \pi_\theta(a^* \mid s)\]
That's behavioural cloning, and it is supervised learning. Not an analogy, the same code. Then from outcomes:
\[\max_\theta\ \mathbb{E}_{\pi_\theta}\Big[ \sum_{t} \gamma^t\, r_t \Big]\]
with \(r_t\) from whether the conversation resolved, escalated, or ended well. Only once the move is chosen does generation run, \(p(\text{words} \mid a, s)\), the seq2seq model one more time.
Example. State: the customer asked a question, nobody has answered, two seconds have passed, the question was addressed to the machine. The policy picks speak. Same question, but the customer turned to a colleague while asking it: wait. Half-heard, with the order number missing: clarify. The training pairs \((s, a^*)\) come from transcripts of human agents in the same spots, and \(r_t\) is \(+1\) if the conversation ends resolved. Only after speak is chosen does the generator write the sentence.
The policy picks the move. Generation only picks the words.
2.11 What you learn from it
Three layers of learning.
- From one conversation: the outcome, the stage path, where it stalled, who moved it.
- Across many: the stage transition grammar \(\phi\), the moves that precede resolution versus escalation, how long each stage takes, who plays which role.
- For a machine participant: the policy \(\pi\). Imitation first, then outcomes. That policy is the interaction layer of the pyramid. Generation is its last step.
In agent terms: disentanglement and tagging are perception, the tracker is state estimation, intent is part of the state, the policy is planning, and generation is execution.
perception → state estimation → planning → execution = disentangle and tag → track → choose a move → say it
The sentence is the last step. The policy that chose it is the product.
3. Other kinds of conversation
Everything in section 2 assumes one style of conversation: live, task-shaped, between people who want the same thing. That's the easy one. Conversations vary on about six dimensions, and each one changes what the tracker has to hold.
Two older fields cover most of this table.
The dimension that matters most for the pyramid is participation role. Almost all conversational AI assumes it's the addressee. A machine in a home or a room is usually a bystander, and knowing which it is right now is the engagement layer.
4. Where we are at the end of 2026
What I see is a pyramid with a finished base and not much above it. Dialogue fell out of language models. The speech community is filling in voice and timing. Almost nobody is working on engagement, presence or agency as product.
4.1 Words and voice
Two layers live here, and it's worth keeping them apart. Dialogue is the organisation of meaning across turns: answering, clarifying, referring back, correcting. Vocal communication is everything the sound carries that isn't a sentence. A laugh, a gasp, a questioning "hmm?" all communicate without forming one.
The words are solved. The voice is close. Speech-to-speech models carry prosody and affect, and 2026 brought a wave of full-duplex work: models that listen while they speak, decide whether to yield or push through an interruption, and call tools without dropping the turn. Most shipped products are still half-duplex. The machine listens or talks, not both. And speaking is still treated as audio to generate rather than an action with a goal, which is what Moore and Nicolao's synthesiser did when it changed how it spoke based on whether it was being understood.
This is also where the habitability gap is easiest to see. The models can say anything, so users can't tell what they can do, and fall back to short commands.
The words are done and the voice is close. Speaking as an action with a goal is not.
4.2 Seeing, timing, acting together
Multimodality means combining evidence, not adding sensors. A robot can have a camera, a microphone and a touch sensor and treat them as three unrelated streams. The useful thing is "that one" + a pointing gesture + the objects in view → the intended object. Vision-language models give the input half of this. The robot sees the scene and the gesture. The output half is empty. No shipped robot gestures, looks or touches as part of talking.
Timing is half done. Turn-taking, overlap and backchannels are what the full-duplex work is about. Space isn't started: approaching without crowding, standing where a listener would stand, moving with a person.
Joint action is not the same as coordination. Two people can avoid bumping into each other without working together. Two people carrying a table are working toward one outcome, and conversation is the second kind. Robots that do tasks exist; the landscape essay covers them. A robot doing a task with a person, each adjusting to the other's hands and pace, is still papers.
The machine sees and it can take a turn. It does not yet act with you.
4.3 Engagement, presence, agency
Presence is not engagement. Someone is standing next to you at a conference. You know they're there. You are not in a conversation with them. For a robot that's the gap between "there is a person here" and "this person is addressing me."
Engagement still mostly means a wake word. Knowing whether you're the one being addressed, in a room with several people in it, is a 2026 benchmark problem. Presence became physical this year. 1X's NEO, the first consumer humanoid sold for the home, ships early-access units in the United States in 2026, voice-first, at $20,000 or $499 a month. Its makers say it recognises when you're speaking to it. Part of its presence is a remote operator, which raises the question of whose presence it is.
Agency is not consciousness, and I got it wrong the first time through because the word covers two things. The narrow reading is goal-directed action with feedback: choose an action, see what it did, adjust. By that reading agency arrived in 2026. A tool-using agent does exactly this, and a robot running a learned policy does it in the world, less reliably. Moore means something more. His architecture rests on perceptual control theory, where an agent acts to keep what it perceives at reference values, and the references come from inside, from needs. By that reading agency is missing. Every 2026 agent's goal arrives in a prompt and dies with the episode. Nothing persists and nothing is its own.
The first kind is what makes an agent useful. The second is what makes it a conversation partner, because a partner has to want something from the exchange. The machines can act. They act only when asked, and a thing that only acts when asked has nothing to say.
The machines can talk. They are not yet there with you, and they have no reason of their own to speak.