Question

Can AI tell who said what in a meeting?

Yes, and the feature has a name: speaker diarization. The AI splits a recording into turns, groups the turns that sound like the same voice, and labels them — Speaker 1, Speaker 2, Speaker 3. What it does not do is know who those people are. Names come from you, or from the conversation itself, and the difference between a transcript that labels speakers well and one that doesn't is usually decided in the first thirty seconds of the recording, not by which app you bought.

Updated August 2026

The short answer

Separating voices: yes, and it works well in normal conditions. Two or three people in a quiet room, one talking at a time, microphone on the table — you'll get clean, consistent speaker turns.

Knowing names: not on its own. The model hears a distinct voice, not an identity. Some tools infer names when people introduce themselves or address each other, and on a video call the platform's participant list helps — but the reliable way to get names into a transcript is for the names to be said out loud, or for you to assign them afterwards.

It degrades with group size. Accuracy in this specific area tends to fall as more people join; Granola's speaker identification, for example, is documented as weakening once three or more people are on a call, and the pattern is general rather than specific to one product.

Why it matters more than it sounds: an action item attached to the wrong person is worse than an action item attached to nobody. Speaker labels are what make automatic task assignment trustworthy.

Noter AI records, transcribes and summarizes your meetings — on iPhone, iPad & Android, in 60+ languages.

How speaker labels actually work

Diarization answers "who spoke when", separately from the transcription that answers "what was said". The two run together and are then stitched into the transcript you read.

The model builds a voice fingerprint from acoustic characteristics — pitch, timbre, speaking rhythm — and clusters segments that match. It's comparing voices to each other within your recording, not to any database of people, which is why the default labels are numbers.

Names get attached in one of three ways. You assign them after the fact. The conversation supplies them ("Thanks, Miriam" — some tools pick this up, and none do it reliably every time). Or, on a video call, the platform identifies which participant's audio stream is active, which is why virtual meetings often produce better-named speakers than in-person recordings do.

The count is inferred, not known. The model works out how many distinct voices are present from the audio. A person who says two words in a ninety-minute meeting may be merged into someone else; a person whose voice changes noticeably — leaning away, turning their head, phone versus room mic — may be split into two.

When it goes wrong

  • People talking over each other. Overlapping speech is the hardest case in the whole field. When two voices share the same moment of audio, the labels are being guessed as much as computed.
  • One microphone, a big table. Someone at the far end arrives quieter and more reverberant than the person next to the phone, and distance changes a voice's acoustic profile enough to confuse the clustering.
  • Similar voices. Two colleagues of the same age, gender and regional accent are a genuinely hard problem, and they're common on a team.
  • Someone joins late or speaks only once or twice. Not enough audio to form a stable fingerprint.
  • A speakerphone in the room. Remote participants arriving through a room speaker often end up merged into a single "speaker", because acoustically that's exactly what they are.
  • Rapid back-and-forth. Short interjections — "yep", "agreed", "no" — get attached to whoever was talking around them more often than in slower exchanges.

How to get better speaker labels

  1. 1Do a round of names at the start. "Miriam, Tom, Chen, and I'm Sara." Ten seconds, said clearly, and it fixes both the labelling and the spelling of names for the rest of the transcript.
  2. 2Put the phone flat in the middle of the table, screen up, nothing covering the microphone. Equal distance beats close-to-you every time for speaker separation.
  3. 3In a larger room, favour the quiet end. The loud voice at the other end of the table will survive the distance; the quiet one won't.
  4. 4Ask for one speaker at a time during the part of the meeting you'll actually want a clean record of. Not the whole hour — just the decisions.
  5. 5Say names when you assign work. "Tom, can you take the pricing sheet?" gives the AI both the owner and the task in one sentence, which is what makes action-item assignment reliable.
  6. 6Avoid a speakerphone daisy chain. If remote people are dialling into a room speaker, they'll arrive as one voice — better to have them join the meeting properly so a bot can capture their streams.
  7. 7Fix the first mislabel, not the twentieth. Correcting a name early in the transcript is faster than untangling it later, and in Noter AI you can edit the transcript and have the AI re-analyse it.

Why this decides whether action items are usable

Speaker labels look like a formatting nicety. They're actually the foundation of the only output that changes what happens after a meeting.

Extraction with the wrong owner is worse than no extraction. A task list where the assignments are 80% right doesn't save anyone time — it creates a round of "I don't think that was mine", and after two of those people stop trusting the list.

Accuracy improves markedly when names are said out loud at the point work is assigned. That single habit does more for action-item quality than any feature comparison, across every tool in this category.

Verification needs to be one tap. The practical answer to a disputed attribution isn't better AI, it's being able to jump straight to the moment and listen. A transcript with timestamps and synced playback turns a disagreement into five seconds of audio.

Tools vary a lot here. Noter AI, Fireflies, Fathom and tl;dv all attach owners when the conversation makes it clear, and pick up deadlines including relative ones like "end of week". Otter and Granola do it less consistently, and Notta barely at all.

What Noter AI gives you

Speaker-labelled transcripts on every note, with timestamps, from in-person recordings as well as from the bot on a Zoom, Teams, Meet or Webex call. In-person meetings get the same treatment as video calls, which is not true of most of this category — Fireflies and Fathom can't capture a conversation in a room at all.

Synced playback is the verification tool. Tap any line in the transcript and the audio jumps to that moment, so checking who actually said something takes one tap rather than a scrub through a recording. This matters more than any claim about diarization quality, because it's what turns a probabilistic label into something you can confirm.

Editable transcripts. Correct a name or a mislabelled turn, and the AI can re-analyse so the summary and action items reflect the fix.

Action items with the person responsible attached, extracted automatically on every note rather than behind a button or metered by credits.

The honest limits. Speaker separation degrades with crowded rooms, overlapping speech and similar voices — that's true here and everywhere else, and no app on the market has solved it. The fastest improvement available to you is a round of names at the start and a phone in the middle of the table. Test it on roughly five minutes of a real meeting with your actual group before subscribing.

Frequently asked questions

Can AI tell who said what in a meeting recording?

Yes — it's called speaker diarization. The AI separates the audio into turns, clusters turns that share a voice, and labels them as separate speakers. It works well with a few people, clear audio and one person talking at a time, and degrades with crowded rooms, overlapping speech and similar-sounding voices.

Does AI know people's actual names, or just Speaker 1 and Speaker 2?

By default it identifies distinct voices, not identities, so labels start as numbers. Names arrive when people introduce themselves or address each other by name, when a video-call platform identifies the active participant, or when you assign them yourself afterwards. A ten-second round of names at the start of an in-person recording is the single most effective thing you can do.

How accurate is speaker identification with a lot of people?

It falls off as the group grows. Two or three voices in a quiet room is close to reliable; six people around a large table with one phone at one end is not. Granola's speaker identification is documented as weakening at three or more, and the pattern is general. Placing the microphone centrally and asking for one speaker at a time during important stretches helps more than switching apps.

Why does the transcript merge two people into one speaker?

Usually because they sound similar, because one of them speaks very little, or because they're both arriving through the same channel — remote participants dialling into a room speakerphone are acoustically one voice. Distance also changes a voice's profile enough to split one person into two, which is the same problem in the opposite direction.

Can I correct speaker labels afterwards?

In Noter AI you can edit the transcript and have the AI re-analyse it, so a corrected name flows through to the summary and action items. Correcting an error early in the transcript is faster than fixing every later instance, and it's worth doing before you circulate a note where the attributions matter.

Do speaker labels work for in-person meetings, not just video calls?

In tools that can record in person at all, yes. Noter AI produces the same speaker-labelled transcript from a phone recording of a meeting in a room as it does from a bot on a video call. Video calls do have an advantage — the platform knows which participant is speaking — so for in-person recordings, saying names out loud early matters more.

Does getting speakers wrong affect the action items?

Yes, and it's the main reason speaker labels matter. A task assigned to the wrong person creates a correction round and, after a couple of those, people stop trusting the list. Saying names out loud when work is assigned — "Tom, can you take the pricing sheet?" — measurably improves assignment accuracy in every tool that extracts action items.

Let Noter AI take your meeting notes

Record, transcribe, and summarize meetings on iPhone, iPad & Android — or send a bot to Zoom, Teams, Meet, or Webex. In 60+ languages.

Related reading