Sally - AI Meeting Assistant

JULY 2026

Meeting Room Acoustics: Why Your Transcript Fails Because of the Room, Not the AI

The same tool produces clean minutes from a video call and fragments from a conference room. Three acoustic effects that every room brings with it explain why.

Microphone at the center of a room with range rings, speakers further away shown fainter

Friday, 10 a.m., large meeting room with a glass wall facing the courtyard. Seven people, two hours of quarterly planning. This time you did everything right and started the recording. That evening you open the transcript and read sentences like "and then we need to sort out the [inaudible] as well". The moment the budget was agreed is exactly one of those gaps.

The obvious explanation would be that the AI is bad. It is not. Modern transcription models reach accuracy on clean audio that was unthinkable five years ago. The problem happens earlier, in the moment sound travels across a room before it reaches a microphone. This article explains the three acoustic effects that destroy information along the way, why better models cannot heal them, and which approach actually solves the problem.

What really happens acoustically in a meeting room

For speech recording, a meeting room is one of the least friendly environments there is. Three effects act at the same time, and they reinforce each other.

Distance works as a square, not a straight line

The common intuition says that sitting twice as far away means being recorded half as loudly. That is too optimistic. In a free field, sound intensity falls with the square of the distance, which means roughly 6 dB of level lost with every doubling of the gap. Someone speaking 30 centimetres from the device lands in the recording at full level. Someone at the far end of a table 2.4 metres away, three doublings out, arrives about 18 dB quieter.

What matters is not loudness itself but the ratio between the wanted signal and background noise. Once a voice sinks close to the noise floor of ventilation, laptop fans and traffic outside, the model loses the contrast it uses to tell sounds apart. This is the threshold where a transcript tips from usable to useless, and it sits closer than most people expect.

Reverberation smears the consonants

The underrated problem is reverberation. Every hard surface in the room throws sound back: glass wall, table top, whiteboard, bare ceiling. So the microphone captures not only the voice but the same voice again a few milliseconds later, several times over, from different directions.

For human hearing this is barely an issue, because the brain separates direct sound from reflections. For a transcription model it is, because reflections cover exactly the short, low-energy sounds that distinguish one word from another. The plosives t, k and p suffer most, along with sibilants. "Deadline" becomes "eadline", "packing" becomes "pa'ing". A modern meeting room with a lot of glass and very little textile is acoustically close to the worst case.

Background noise nobody notices during the meeting

While talking, everyone filters out the surroundings. A microphone cannot. It records the air conditioning, the projector fan, cups on saucers, a chair scraping the floor and the keyboard of the person taking notes. These sounds sit in the same frequency range as speech, so they cannot simply be filtered out.

Constant noise like ventilation is the most damaging, because it lifts the noise floor permanently. That moves the distance at which a voice disappears into the noise closer for the entire recording. If you can change only one thing in a meeting room, switch off the ventilation.

Why a better model does not rescue this

The hope that the next model update will fix it is understandable and unfortunately wrong. It is worth understanding why.

The model guesses, and guessing is systematic

A transcription model assigns the most probable sequence of words to an audio signal. When the acoustic information is missing, it fills the gap with whatever is linguistically plausible. The result is more dangerous than a gap: you get grammatically correct sentences that are factually wrong. One price becomes a different price, a "not" disappears, a name turns into the nearest similar word in the vocabulary.

With bad handwriting you notice that you are guessing. With a fluent transcript you do not. That is why recording quality is not a comfort question but a question of how far your documentation can be trusted.

Speaker separation without an anchor

The second loss concerns attribution. In online meetings each participant supplies a separate audio track, so attribution is technically unambiguous. In a room everything lands on one track and the software has to work out who is speaking from sound alone. With similar voices, comparable pitch and a reverberant room, that produces mix-ups.

This is the point where minutes lose their value. A transcript where the words are right but assigned to the wrong person is worse than no transcript, because it creates trust it has not earned. Why attribution is so hard in a room and what makes it reliable is covered in the article on speaker separation with multiple people.

What you can improve about room acoustics, and where the ceiling is

Before any technology: some of this can be fixed with zero budget, and the effect is bigger than expected.

What helps immediately and costs nothing

Four measures measurably improve recordings. First, seating: a tight circle around the device beats a spread-out boardroom layout, because it halves the maximum distance. Second, placement: put the device openly in the middle of the table, not on paper, not in a corner, not beside the projector. Third, switch off ventilation for the duration. Fourth, bring soft surfaces into the room, which in practice usually means closing curtains and the door, and not booking the room with four glass walls when another one is free.

Where hardware stops helping

A good conference microphone with an omnidirectional pattern, or an array with direction detection, captures more wanted signal than a phone lying flat on a table. That advantage is real. It does not change the principle: one device records from one position, and that position is the right one for at most two people.

This is why hardware scales so badly against this particular problem. A microphone that costs twice as much does not double the range, because the level drop stays quadratic and the reverberation stays in the room. Beyond a certain group size, a room microphone cannot meaningfully improve the result, whatever the budget.

The way out: close microphones instead of one room microphone

All three effects above share one root cause, which is distance. Reverberation matters because the direct sound arrives weak. Background noise matters because the wanted signal is small. Overlapping voices become a problem because they land on one track. Remove the distance and all three go away at once.

That is exactly the approach behind transcription of in-person meetings with Sally: instead of hunting for a better microphone for the room, several phones in the room are linked into one shared recording. Each phone becomes the close microphone for the person in front of it. That is the actual technical core, and it is worth a closer look.

How linking works during the meeting

One person starts the recording in the Sally app. The other participants link their own phone to that same recording, so in the end several devices jointly record one single meeting rather than producing several separate captures. Each device sits visibly in front of its owner on the table, which also makes consent transparent.

The effort is a few seconds per person, and it needs no setup in the room, no pairing with an installed system and no rights on the conference hardware. Teams with fixed meeting rooms can additionally connect the room calendar, so Sally recognises by itself when a meeting starts there.

What technically happens to the individual tracks

What happens after the meeting is the decisive part. The devices do not deliver one mixed room recording but one separate audio track each, captured from short range. Sally then merges those tracks into one clean audio track while keeping the information about which voice came from which device.

That removes all three opponents from this article at once. Distance is gone, because every voice is captured from a few centimetres. Reverberation loses weight, because on each track the direct sound is far stronger than the reflections. And overlapping voices stay separable, because they physically sit on different tracks instead of adding up into one signal. The result is speaker-separated and reaches up to 98.8 % accuracy according to Sally, without anyone having to buy, charge or carry hardware.

Why quality rises with every device

One practical point: a single phone is entirely enough to document a meeting. It then works under the same physical conditions as any other single device, covering the case where otherwise nothing would be documented at all.

Every additional linked device improves the result, because one more voice is captured from close range and gets its own channel. So the approach behaves exactly the opposite way to a room microphone: a microphone in the middle of the table gets worse with every additional person, while linked phones get better with every additional person. The advantage therefore grows precisely where hardware hits its ceiling, which is in large groups.

Which recording situation leads to which quality

SituationTypical distanceCritical factorExpected transcript quality
Two people, device between them30 to 60 cmbarely a constraintvery good
Four people, small tableup to 1.2 mreverberationgood, occasional gaps
Eight people, long tableup to 3 mdistance plus overlappatchy, speakers often wrong
Large room with glass surfacesover 3 mreverberation dominatesmostly unusable
Several linked phonesunder 60 cm per personnone of the three effects applyvery good, speaker-separated

Whatever the technology, recording a conversation needs a legal basis. In Germany, for example, Section 201 of the Criminal Code makes secretly recording the spoken word a criminal offence, so recordings are only permitted with the consent of everyone involved, and comparable two-party consent rules exist in several other jurisdictions. In practice this means announcing the recording at the start and obtaining agreement.

Phones lying visibly on the table are an advantage here, because they make the documentation transparent, unlike a discreet device in a pocket. Sally stores and processes exclusively in Germany, GDPR-compliant and without third-country transfer. What that means in detail is on the page about GDPR and security.

Conclusion

When a transcript from a meeting room disappoints, the cause is almost never the model and almost always the room. Distance cuts level quadratically, reflections smear consonants, constant noise raises the noise floor. No software can recover the lost information from that source material, it can only guess plausibly.

Seating, placement and switched-off ventilation get you a real improvement at no cost. The actual solution, though, is a change of principle: not a better microphone for the whole room, but one microphone per voice. The devices for that are already on the table.

Try it at your next real meeting: test Sally free for 30 days, or first read how transcribing in-person meetings works overall.

FAQ

Lorenz Zwicknagl

Lorenz Zwicknagl

Marketing

Meetings should be a means of solving problems, not another waste of time. Artificial intelligence can help make them more efficient by summarizing discussions, highlighting key points, and clearly defining tasks. This creates more room for decisions instead of repetitions.

Learn more about the author