Sally - AI Meeting Assistant

JULY 2026

Speaker Separation With Multiple People: Why Minutes Quote the Wrong Person

A transcript with the right words and the wrong speaker is worse than no transcript. Why attribution breaks down in a room, and how it becomes unambiguous.

Four speakers on the left whose voices converge as separate tracks into one transcript

Two weeks after the kick-off, a discussion escalates. The minutes say the project lead committed to the migration date. He says that was not him, that was the colleague from IT. Both are right: the words in the transcript are correct, the name above them is wrong.

This kind of error is more insidious than a gap in the text, because it is invisible. A transcript containing "[inaudible]" warns you. A transcript with wrong speaker attribution reads smoothly and convincingly. This article explains why attribution is a separate technical problem, why it systematically breaks down in a meeting room, and what makes it unambiguous.

What speaker separation technically is

The first step to understanding this is a distinction that often blurs in practice: transcription and speaker attribution are two different jobs, solved by different methods.

Diarisation answers a different question

Transcription answers which words were spoken. Speaker separation, technically diarisation, answers who spoke when. It splits the recording into segments and groups those segments by person, without needing to know who the people are. That is why results start out as "Speaker 1" and "Speaker 2" until someone assigns names.

What matters is that both processes can be good or bad independently of each other. A transcript can be word-perfect and still completely misattributed. That specific combination produces the most dangerous minutes, because the linguistic quality creates trust the attribution has not earned.

How software tells voices apart

Diarisation works from acoustic features: the fundamental frequency of a voice, timbre, speech rhythm, the characteristic resonances of a person's vocal tract. From these features it builds a kind of fingerprint per speech segment, and similar fingerprints are assigned to the same person.

That works well as long as the fingerprints sit far apart and stay stable. In a meeting room both conditions are violated, for three reasons.

Why attribution systematically breaks down in a room

It helps to see that this is not a question of software quality. The causes sit in the signal.

Similar voices really are close together

Two men of comparable height and similar age often have a very similar fundamental frequency and similar resonances. The same applies to two women with a similar register. Add a similar speaking pace and the same regional accent, and the acoustic fingerprints sit so close that they overlap.

At that point the software is not making a decision, it is making a probability estimate, and in a share of cases that estimate is wrong. In a group of eight, the number of possible confusion pairs rises sharply, which is why large groups get harder more than proportionally.

Simultaneous speech is mathematically unsolvable

The hardest case is overlap. When two people talk at once, their sound waves add up on the way to the microphone into a single signal. Recovering two original sources from that one signal is an underdetermined problem: two unknowns, one equation. There is no unique solution, only more or less plausible approximations.

In practice, the overlapping passage gets assigned to one person, usually the louder or closer one, and the other contribution disappears from the minutes. This happens most often in lively discussions, exactly where the important objections are raised.

Reverberation makes a voice unlike itself

The third reason is subtle. Reflections in the room change timbre depending on where someone sits and which way they are facing. When a person turns to the whiteboard, they sound different from how they sounded two minutes earlier at the table.

For diarisation that means a single person's fingerprint drifts across the recording. The consequence is a second type of error: instead of confusing two people, one person gets split into two speakers. Minutes containing "Speaker 3" and "Speaker 6" who are obviously the same person come about exactly this way.

Why online meetings do not have this problem

The most striking everyday observation is that the same tool produces clean attribution on video calls. The reason is not a better algorithm but a better starting position.

In Google Meet, Zoom, Microsoft Teams and Webex, every person sits at their own microphone and supplies a separate audio track with an unambiguous participant identifier. The software does not have to guess who is speaking at all, it reads it from the channel. Overlaps are harmless, because both voices sit on different tracks and both can be transcribed in full.

That names the actual insight: reliable speaker separation is not a question of computing power but a question of channels. Anyone who wants the same quality in a room needs the same structure there.

What makes attribution unambiguous in a meeting room

This is exactly where the approach behind transcription of in-person meetings with Sally comes in. Several phones in the room are linked into one shared recording, and each phone becomes the close microphone for the person in front of it.

One device per voice is a hard anchor

This brings the channel back as a source of information. Who is speaking no longer follows from an acoustic estimate but from the device on which the voice arrives loud and close. All three error sources above lose their basis: similar voices are harmless because they sit on different tracks. Overlaps stay separable. And the drifting fingerprint no longer matters, because attribution does not depend on sound.

Sally then merges the tracks into one clean audio track. The result is speaker-separated and reaches up to 98.8 % accuracy according to Sally, with no extra hardware.

What this changes in the minutes

The practical difference only shows up after the meeting. When attribution is right, tasks can be assigned automatically to the correct person, decisions come with traceable reasoning, and the question of who actually committed to something becomes a search rather than an argument. That is the point at which a recording turns into documentation you can rely on.

Methods compared

Recording setupAttribution based onOverlapTypical errors
One room microphone, monosound, estimatednot separablemix-ups, one person split in two
Conference microphone with arraysound plus directionpartly separableerrors for people sitting close together
Wearable hardware recordersound, estimatednot separablewearer dominates, group indistinct
Online meeting, track per personparticipant channelfully separablepractically none
Several linked phonesdevice channelfully separablepractically none

Honest limits

Fair is fair: several phones do not make a transcript magically perfect either. When three people share one device, acoustic estimation is still needed for those three. Anyone speaking very quietly or covering their mouth is captured less well. And if nobody but you links a phone, Sally works with one track, under exactly the same physical conditions as any other single device.

The difference is that quality here grows with participation rather than with a hardware budget. Where a conference microphone hits a hard ceiling at eight people, the result improves with every additional linked phone. How the two approaches compare directly is covered in conference microphone or several phones.

Conclusion

Speaker separation is not a side aspect of transcription but the property that turns a wall of words into usable documentation. On a single audio track it stays an estimate, and that estimate reliably fails exactly where it matters: with similar voices, in lively discussions and in reverberant rooms.

It becomes reliable when every voice has its own channel. Online meetings have that by nature, and in a meeting room the same structure can be built from the phones that are already on the table.

Try it at your next meeting with several people: test Sally free for 30 days, or read the overview on transcribing in-person meetings first.

FAQ

Lorenz Zwicknagl

Lorenz Zwicknagl

Marketing

Meetings should be a means of solving problems, not another waste of time. Artificial intelligence can help make them more efficient by summarizing discussions, highlighting key points, and clearly defining tasks. This creates more room for decisions instead of repetitions.

Learn more about the author