Florence D. Jiang

Towards Intelligent Communicative Machines

23 min read

I worked on speech recognition for many years, and it's an almost solved problem. That is an exciting thing to be able to say. It is also a strange one, because it feels like the field is finished and the problem I spent years on is gone.

Then at Interspeech 2025 I ran into a slide from Roger Moore. It was a pyramid with eight layers. Spoken dialogue, the part I had been working on, was the base. Above it were seven more: knowing someone is there, knowing they are talking to you, timing, doing things together, and at the top, having a reason to speak at all. Each layer builds on the one below.

The problem was never recognition. It was understanding, and that one is always there. Recognition was the first step up a pyramid, and the rest of it is the road to a more intelligent communicative machine.

1. The pyramid

Read it from the bottom. It goes from being able to talk to having a reason to.

Moore's pyramid: eight prerequisites for spoken-language interactionAgency at the apex, spoken-language dialogue at the base.

After R. K. Moore, Interspeech 2025 keynote. Labels from the slide.

Layer, top to bottomWhat it meansWhat the robot does
1. AgencyI can act toward a goal: choose an action, see what it did, adjustIt can try to get you a cup of water, rather than describe doing so
2. PresenceWe are available to one another. It notices a person and lets the person notice itIt sees you approach and turns toward you
3. EngagementWe are actually interacting. Being nearby is different from being in an exchangeIt can tell you asking for help from you talking to someone beside it
4. Spatio-temporal coordinationWe coordinate where and when things happenIt approaches without crowding you, answers at the right moment, pauses when you interrupt
5. Joint actionWe work toward one outcome, and each one's actions fit the other'sYou hold out your hand; it brings the cup within reach and manages the handover
6. Multimodal interactionWe combine signals. Words, gaze, gesture and movement count togetherYou say "that cup" while pointing, and it uses both
7. Vocal communicationWe communicate through sound, not just written words read aloudIt hears your "hmm?" as a request to clarify and changes how it speaks
8. Spoken-language dialogueWe use spoken language to build shared understanding across turnsYou say "the other one"; it connects that to the last exchange and checks what you mean

A machine with no goals has no reason to speak. A machine that isn't present can't engage. A machine that can't coordinate in time can't take a turn. A machine that can't act with you has nothing to talk about. The words are necessary, and they are the smallest part.

Moore has a second idea that explains what the base feels like on its own. He calls it the habitability gap. As a system gets more flexible, usability drops, because people can no longer tell what it can do. He reads it as the uncanny valley, and says you only avoid it when "the visual, vocal, behavioural and cognitive affordances of an artefact" match one another. A perfect voice on a machine that can't engage is a mismatch, and people feel it before they can say why.

The habitability gap is what a missing layer feels like from the outside.

2. The ML behind the pyramid

Here's the part that made the pyramid feel like home. Every layer is an ML problem I already knew, applied to a bigger unit. The lineage is token classification, sequence tagging, sequence to sequence, translation, language model, masked language model, causal transformer. Token-level tasks label tokens. Conversation-level tasks label utterances, turns and threads. The models changed. The formulations didn't.

The map first, then each method on its own with the one equation that defines it. Throughout, \(x\) is the input sequence, \(y\) the output, \(h\) a hidden vector the model computes, and \(\theta\) the parameters.

ProblemUnitFormulationClassic methodNow
Who is talking to whom (disentanglement)Message → threadOnline clusteringReply-to scoring between each new message and the last N, then link or open a new threadTransformer embedding of (speaker, time, text), same linking rule
Dialogue actUtterance → moveClassificationSVM or CRF over utterance features: question, statement, backchannel, agreementFine-tuned encoder, or zero-shot with a schema
StageUtterance → stageSequence tagging over utterancesHMM or CRF: stages follow a grammar, so tag with transition constraintsHierarchical transformer: encode each utterance, then a sequence model over utterances
Stage boundariesPosition → boundary or notToken classification at utterance levelTextTiling-style segmentationSame, with learned embeddings
StateConversation → structured statestate_t = f(state_t−1, utterance_t)Slot filling: a belief state over slotsSeq2seq that emits a state diff as JSON
IntentUtterance → intent, person → goalClassification, then inferenceIntent classifierClassifier for surface intent; inverse planning for the goal
ResponseState → action → wordsPolicy, then generationHand-written policyLearned policy, LLM generation

Conversation understanding is token classification with a bigger token.

2.1 Token classification: Label every token

The simplest formulation. Each position gets its own label, and the labels don't talk to each other. Encode the sequence, put a softmax over labels on every position:

\[p(y_i \mid x) = \mathrm{softmax}(W h_i + b), \qquad \mathcal{L} = -\sum_{i} \log p(y_i \mid x)\]

Read it left to right: a vector for each token, a distribution over labels from that vector, and a loss that makes the right label likely. That's named-entity recognition. At the utterance level the same head answers "does a new stage start here?", with \(y_j \in \{\text{boundary}, \text{no boundary}\}\).

Example. x = [I, love, New, York], y = [O, O, B-LOC, I-LOC]. Four tokens, four labels: O for outside any entity, B-LOC for the start of a place name, I-LOC for its continuation. The model scores York as I-LOC from \(h_{\text{York}}\) alone. It never checks what it gave New.

Token classification is a softmax on every position, trained independently.

2.2 Sequence tagging: Labels that depend on each other

Labels are rarely independent. A question is usually followed by an answer. A conversation that has reached its closing doesn't jump back to the opening. Sequence tagging adds a term for the transition between neighbouring labels and scores the whole label sequence at once. This is the conditional random field:

\[p(y \mid x) = \frac{1}{Z(x)} \exp\Big( \sum_{i} \psi(y_i, x) + \sum_{i} \phi(y_{i-1}, y_i) \Big)\]

\(\psi\) says how well label \(y_i\) fits position \(i\). \(\phi\) says how plausible the step from the previous label is. \(Z(x)\) normalises over every possible label sequence. Decoding is Viterbi, which finds the best path in \(O(n \cdot |Y|^2)\) instead of trying them all. The hidden Markov model is the generative cousin, \(p(x, y) = \prod_t p(y_t \mid y_{t-1})\, p(x_t \mid y_t)\), and it's what stage detection ran on before transformers: a handful of stages, a transition matrix, an emission model over utterance features.

Example. Same x. The path [O, O, O, I-LOC] is impossible, because an entity can't continue without starting, so \(\phi(\text{O}, \text{I-LOC}) = -\infty\) and Viterbi never returns it. At the utterance level: x = ["Where is the exit?", "Down the hall.", "Thanks."], y = [question, answer, closing]. \(\phi(\text{question}, \text{answer})\) is high and \(\phi(\text{closing}, \text{question})\) is low, so the tagger prefers the path a conversation would actually take.

Sequence tagging scores the whole label path, so a stage can't jump where the grammar forbids.

2.3 Sequence to sequence: One sequence becomes another

When the output is a sequence with its own length, the model generates it one element at a time, each conditioned on the input and on everything it has generated so far:

\[p(y \mid x) = \prod_{t=1}^{T} p(y_t \mid y_{<t}, x)\]

Translation was the first use. The same factorisation turns a conversation into a state diff, a thread into a summary, a state into a reply. The encoder reads \(x\) into vectors and the decoder produces \(y_t\) by attending back to them.

Example. x = I love New York, y = J'aime New York. The decoder emits J' given \(x\), then aime given \(x\) and J', then New, then York, then a stop token. Nothing says four tokens in means four tokens out. Later in the essay the same model takes x = previous state + new utterance and emits y = {"order_id": "A1234"}.

Seq2seq generates the output token by token, conditioned on the input and on itself.

2.4 Language models: Predict the next token

Drop the input sequence and the seq2seq equation becomes a language model:

\[p(x) = \prod_{t=1}^{T} p(x_t \mid x_{<t})\]

That's the causal, left-to-right model, and it's the one that generates. The masked language model trains the other way round. Hide some tokens and predict them from both sides:

\[\mathcal{L}_{\text{MLM}} = -\sum_{i \in M} \log p(x_i \mid x_{\setminus M})\]

It can't generate, but its hidden vectors \(h_i\) have seen the whole sequence, which makes them the better input to a classifier or a tagger. In conversation work the encoders label things and the causal models say things.

Example. Causal: x = I love New and the model puts most of its probability on York. Masked: x = I [MASK] New York and the model predicts love, using I on the left and New York on the right. The vector for [MASK] saw the whole sentence, which is why encoder vectors make better features for labelling.

Causal models generate. Masked models represent. Both are next-token prediction with a different mask.

2.5 The causal transformer: One model for every formulation

All of the above now run on one architecture. Attention lets every position read every earlier one:

\[\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^\top}{\sqrt{d}} + M\right) V\]

\(Q\), \(K\) and \(V\) are linear projections of the hidden vectors. \(M\) is the causal mask, \(-\infty\) above the diagonal, so a token can't see its future. The same hidden state \(h_t\) feeds whatever head you put on it: a softmax for classification, a CRF for tagging, the next-token distribution for generation. And a prompt turns any task into generation, \(p(y \mid \text{instruction}, x)\). That's why one model now does the whole map table.

Example. For x = I love New York, the query from York attends to the keys of I, love, New and itself. For love, the mask hides New and York. The weight from York onto New comes out high, so \(h_{\text{York}}\) carries "second half of a place name." Put a label head on it and you get I-LOC. Put the next-token head on it and you get a guess at what follows. Put Tag the places in: in front and the same model writes the labels out as text.

One hidden state, many heads. The head, or the prompt, picks the formulation.

2.6 Embeddings and online clustering: Who is talking to whom

Several people in one channel, and the first job is to split the stream into conversations. A message gets an embedding that mixes its text with who sent it and when:

\[e_m = f_\theta(\text{text}_m, \text{speaker}_m, t_m)\]

Each new message is scored against the threads still open in a window. Cosine similarity to the thread's running embedding is the baseline. A learned scorer that also sees the time gap and whether this speaker is already in the thread does better:

\[s(m, T) = \sigma\big( w^\top [\, e_m ;\ e_T ;\ \Delta t ;\ \mathbb{1}[\text{speaker}_m \in T] \,] \big)\]

The assignment rule is the online part:

\[T^* = \arg\max_T s(m, T), \qquad \text{join } T^* \text{ if } s(m, T^*) > \tau, \text{ else open a new thread}\]

and the thread embedding updates as a running mean, \(e_T \leftarrow \frac{n e_T + e_m}{n + 1}\). Evaluate it as clustering against hand-labelled threads, with variation of information or exact-match thread counts. In speech the same problem is diarization plus overlap.

Example. A team channel:

  • m1 Ann, 10:00: "Anyone up for lunch?"
  • m2 Bo, 10:00: "The build is broken again"
  • m3 Cy, 10:01: "Sure, noodles?"
  • m4 Ann, 10:01: "Which test?"

x is each message with its speaker and time. y is a thread id. m3 scores high against m1 because it answers an open question, so T1 = {m1, m3}. m4 is the interesting one. By speaker it belongs with m1, since Ann started that thread. By content it answers m2, so it joins T2 = {m2, m4}. Text alone gets it wrong. Speaker alone gets it wrong. Reply structure wins.

Disentanglement is argmax-and-threshold over a learned similarity, one message at a time.

2.7 Hierarchical sequence model: Stage per utterance

Stage is a label on an utterance, and an utterance is itself a sequence, so the model has two levels. Encode each utterance \(j\) into one vector, then run a sequence model over the utterance vectors:

\[u_j = \mathrm{Enc}(x_{j,1}, \dots, x_{j,n_j}), \qquad h_j = \mathrm{Seq}(u_1, \dots, u_j), \qquad p(z_j \mid h_j) = \mathrm{softmax}(W h_j)\]

\(z_j\) is the stage. Put the transition term from 2.2 over the \(z_j\) and it's a CRF at the utterance level. That's where the stage grammar lives: \(\phi(z_{j-1}, z_j) = -\infty\) for transitions the domain forbids. Per-person stage is the same model run on one speaker's utterances, and the gap between a speaker's \(z_j\) and the thread's is a feature in its own right.

Example. A support chat, x = five utterances in order:

  1. "Hi, my order hasn't arrived"
  2. "Order number?"
  3. "A1234"
  4. "I've reshipped it"
  5. "Thanks, bye"

y = [problem, clarify, clarify, resolve, close]. Each \(u_j\) is the encoded utterance, each \(h_j\) has read the ones before it, and \(\phi(\text{close}, \text{problem}) = -\infty\) stops the tagger reopening a closed conversation. Run it per speaker and the customer is still at problem while the agent is at clarify, which is normal, or at resolve, which means the agent moved on too early.

Stage detection is sequence tagging where the token is an utterance.

2.8 State tracking: Update a belief

This is the step I wrote about in Robotics 101 under state estimation, with utterances in place of sensor readings. Utterances are noisy observations. Keep a belief \(b_t(s)\) over states, predict how it drifts, correct it with what was just said:

\[b_t(s) \propto p(u_t \mid s) \sum_{s'} p(s \mid s')\, b_{t-1}(s')\]

\(p(s \mid s')\) is the transition model, the stage grammar again. \(p(u_t \mid s)\) is how likely this utterance is in that state. For a state with many fields the neural version is a recurrence,

\[s_t = g_\theta(s_{t-1}, u_t)\]

and the current version generates only what changed, \(p(\Delta s_t \mid s_{t-1}, u_t)\): the seq2seq model from 2.3, reading the previous state and the utterance and emitting the fields to update. Incremental, because it reads one utterance. Probabilistic, because it keeps \(b_t\) rather than a label. Typed, because \(s\) has fields that code can branch on.

Example. Same chat. After utterance 1 the state is {issue: missing order, order_id: ?, stage: problem} with a belief over stage of {problem: 0.8, clarify: 0.2}. Utterance 2, "Order number?", moves the belief to {clarify: 0.9} and leaves the fields alone. Utterance 3, "A1234", changes one field, so the diff model emits {"order_id": "A1234"} and nothing else. The tracker never re-reads utterances 1 and 2. What it needed was already in \(s_{t-1}\).

State tracking is a Bayes filter where the sensor is the last thing somebody said.

2.9 Intent: Classify the move, infer the goal

Surface intent is classification, \(p(a \mid u)\), with \(a\) from a fixed set of acts. The goal is never observed. It has to be inferred from the whole sequence of moves:

\[p(g \mid u_{1:t}) \propto p(g) \prod_{i=1}^{t} p(u_i \mid g, s_{i-1})\]

Which goal makes the moves so far most likely? The inverse-RL version says the same thing with a reward. Find \(r\) such that the observed moves are near-optimal under the policy it induces:

\[\pi^*(a \mid s;\, r) \propto \exp\big(Q_r(s, a)\big), \qquad r^* = \arg\max_r \sum_i \log \pi^*(a_i \mid s_i;\, r)\]

Social intent is the same machinery with a much wider hypothesis space, so the posterior stays flat. That is the formal way of saying it's always a guess.

Example. u = "Is there any way to get this faster?" Surface intent: request. That's the label. The goal needs the sequence: the customer mentioned Friday twice, asked about express shipping, never mentioned a refund. The posterior comes out around {needs it by Friday: 0.7, wants a refund: 0.1, other: 0.2}. Social intent: "I guess that's fine" is an acceptance with a hedge, and the best estimate is something like {accepts: 0.6, unhappy but conceding: 0.4}. It should stay that uncertain.

Surface intent is a label. Goal is a posterior over what would explain the moves.

2.10 Policy: Choose the next move

Given the state, the machine picks an action from a small set, \(a \in \{\text{speak}, \text{wait}, \text{acknowledge}, \text{clarify}, \text{yield}\}\). First by imitation, from human conversations:

\[\max_\theta \sum_{(s, a^*)} \log \pi_\theta(a^* \mid s)\]

That's behavioural cloning, and it is supervised learning. Not an analogy, the same code. Then from outcomes:

\[\max_\theta\ \mathbb{E}_{\pi_\theta}\Big[ \sum_{t} \gamma^t\, r_t \Big]\]

with \(r_t\) from whether the conversation resolved, escalated, or ended well. Only once the move is chosen does generation run, \(p(\text{words} \mid a, s)\), the seq2seq model one more time.

Example. State: the customer asked a question, nobody has answered, two seconds have passed, the question was addressed to the machine. The policy picks speak. Same question, but the customer turned to a colleague while asking it: wait. Half-heard, with the order number missing: clarify. The training pairs \((s, a^*)\) come from transcripts of human agents in the same spots, and \(r_t\) is \(+1\) if the conversation ends resolved. Only after speak is chosen does the generator write the sentence.

The policy picks the move. Generation only picks the words.

2.11 What you learn from it

Three layers of learning.

  • From one conversation: the outcome, the stage path, where it stalled, who moved it.
  • Across many: the stage transition grammar \(\phi\), the moves that precede resolution versus escalation, how long each stage takes, who plays which role.
  • For a machine participant: the policy \(\pi\). Imitation first, then outcomes. That policy is the interaction layer of the pyramid. Generation is its last step.

In agent terms: disentanglement and tagging are perception, the tracker is state estimation, intent is part of the state, the policy is planning, and generation is execution.

perception → state estimation → planning → execution = disentangle and tag → track → choose a move → say it

The sentence is the last step. The policy that chose it is the product.

3. Other kinds of conversation

Everything in section 2 assumes one style of conversation: live, task-shaped, between people who want the same thing. That's the easy one. Conversations vary on about six dimensions, and each one changes what the tracker has to hold.

StyleWhat's differentWhat the tracker needs that it didn't
Social talk, small talkNo task, no stages. The point is the relationship, not the contentRapport and engagement signals; a "still worth continuing?" estimate instead of a stage
Asynchronous: email, forums, issue threadsGaps of hours or days; threads interleave over weeks; replies quoteTime becomes a weak signal and quoting a strong one; a thread is a reply graph, not a sequence
Spoken rather than writtenOverlap, backchannels, disfluency, repair ("no, the other one"), prosodyDiarization; a turn-taking state (who holds the floor, who is bidding for it); repair as a first-class move
One-to-many: lecture, broadcast, podcastOne speaker; the audience responds through side channels or not at allAudience state (attention, confusion) without utterances
Asymmetric roles: interview, teaching, support, doctor and patientOne party drives, and the stage grammar belongs to the driver. Teaching has its own pattern: initiate, respond, evaluateA role per participant, and whose grammar applies
Adversarial: negotiation, debate, interrogationGoals conflict; people hide intent; silence and stalling are movesIntent modelled as strategic, with deception possible; game-theoretic rather than cooperative inference
Deliberation: decision meetings, brainstormingMany threads converge on one decision; action items are the outputProposal tracking: who agreed, what was decided, who owns what
Emotional support, ventingInformation content is low; the goal is regulation, not resolutionAn affect trajectory; listening as an action; resolution is not the success metric
Group dynamicsDominance, coalitions, floor hogging, people being talked overA social graph over the conversation, not just a thread graph
Long-horizon relationshipsThe conversation spans months across sessionsMemory: what was said before, what changed, what the person already knows
Machine as bystanderThe machine isn't addressed. It listens to people talk to each other and decides whether to step inParticipation roles: speaker, addressee, side participant, bystander. The single most important state for a machine in a room
Embodied, non-verbalGaze, gesture, distance and touch carry the turnThe spatial layers of the pyramid; sign language is a whole modality
Multilingual, cross-culturalCode-switching; tolerance for overlap and silence differs by cultureLanguage id per utterance; culture-specific turn-taking priors
Machine to machineAgents talking to agentsProtocols, not pragmatics. A different essay

Two older fields cover most of this table.

And four formulations the map in section 2 leaves out:

The dimension that matters most for the pyramid is participation role. Almost all conversational AI assumes it's the addressee. A machine in a home or a room is usually a bystander, and knowing which it is right now is the engagement layer.

Most conversational AI assumes it is being spoken to. In a room, it usually isn't.

4. Where we are at the end of 2026

The pyramid at the end of 2026How much of each layer exists, in products. Green is built, grey is not.exists in productsmissingpartly there

Judgement from products and papers through 2026; the evidence is in the text below.

What I see is a pyramid with a finished base and not much above it. Dialogue fell out of language models. The speech community is filling in voice and timing. Almost nobody is working on engagement, presence or agency as product.

4.1 Words and voice

Two layers live here, and it's worth keeping them apart. Dialogue is the organisation of meaning across turns: answering, clarifying, referring back, correcting. Vocal communication is everything the sound carries that isn't a sentence. A laugh, a gasp, a questioning "hmm?" all communicate without forming one.

The words are solved. The voice is close. Speech-to-speech models carry prosody and affect, and 2026 brought a wave of full-duplex work: models that listen while they speak, decide whether to yield or push through an interruption, and call tools without dropping the turn. Most shipped products are still half-duplex. The machine listens or talks, not both. And speaking is still treated as audio to generate rather than an action with a goal, which is what Moore and Nicolao's synthesiser did when it changed how it spoke based on whether it was being understood.

This is also where the habitability gap is easiest to see. The models can say anything, so users can't tell what they can do, and fall back to short commands.

The words are done and the voice is close. Speaking as an action with a goal is not.

4.2 Seeing, timing, acting together

Multimodality means combining evidence, not adding sensors. A robot can have a camera, a microphone and a touch sensor and treat them as three unrelated streams. The useful thing is "that one" + a pointing gesture + the objects in view → the intended object. Vision-language models give the input half of this. The robot sees the scene and the gesture. The output half is empty. No shipped robot gestures, looks or touches as part of talking.

Timing is half done. Turn-taking, overlap and backchannels are what the full-duplex work is about. Space isn't started: approaching without crowding, standing where a listener would stand, moving with a person.

Joint action is not the same as coordination. Two people can avoid bumping into each other without working together. Two people carrying a table are working toward one outcome, and conversation is the second kind. Robots that do tasks exist; the landscape essay covers them. A robot doing a task with a person, each adjusting to the other's hands and pace, is still papers.

The machine sees and it can take a turn. It does not yet act with you.

4.3 Engagement, presence, agency

Presence is not engagement. Someone is standing next to you at a conference. You know they're there. You are not in a conversation with them. For a robot that's the gap between "there is a person here" and "this person is addressing me."

Engagement still mostly means a wake word. Knowing whether you're the one being addressed, in a room with several people in it, is a 2026 benchmark problem. Presence became physical this year. 1X's NEO, the first consumer humanoid sold for the home, ships early-access units in the United States in 2026, voice-first, at $20,000 or $499 a month. Its makers say it recognises when you're speaking to it. Part of its presence is a remote operator, which raises the question of whose presence it is.

Agency is not consciousness, and I got it wrong the first time through because the word covers two things. The narrow reading is goal-directed action with feedback: choose an action, see what it did, adjust. By that reading agency arrived in 2026. A tool-using agent does exactly this, and a robot running a learned policy does it in the world, less reliably. Moore means something more. His architecture rests on perceptual control theory, where an agent acts to keep what it perceives at reference values, and the references come from inside, from needs. By that reading agency is missing. Every 2026 agent's goal arrives in a prompt and dies with the episode. Nothing persists and nothing is its own.

The first kind is what makes an agent useful. The second is what makes it a conversation partner, because a partner has to want something from the exchange. The machines can act. They act only when asked, and a thing that only acts when asked has nothing to say.

The machines can talk. They are not yet there with you, and they have no reason of their own to speak.

Sources