The Hybrid Meeting Problem

Hybrid meeting audio and speaker identification already have solutions. The bigger opportunity is keeping the recording alongside approved decisions.

Published

Tags: ,

Eight people in a room. Five dialled in. One agenda, but two different experiences. The room has eye contact, quick asides and its own rhythm. Everyone else has to find a gap in the call.

Three familiar failure modes:

Eight people share a room and local cues; five remote participants reach them through one room tile and shared audio. Edit this diagram on Excalidraw

Start with one computer per person

Everyone joins the call individually, with their own computer, camera and headset. Even the people in the office. Each person has a named presence, access to chat and hand-raising, and can join a mixed breakout. GitLab recommends this, preferably with people moving into separate workspaces.

It’s a clear solution, and often an impractical one. An office may not have eight quiet spaces available. Headsets around one table still pick up neighbouring voices, and people hear each other both directly and through the call. It also gives up some of the reason for gathering in a room.

There is already a useful alternative: Google Meet’s adaptive audio coordinates the microphones and speakers of several laptops in the same room to prevent echo. It requires an eligible Workspace subscription and doesn’t work alongside Meet hardware devices. One computer per person needn’t always mean separate rooms or headsets. However, Meet merges the room’s audio and highlights its participants collectively; this isn’t a promise of separate recordings or named speaker attribution.

One computer per person gives everyone access to a shared call and agenda, but needs quiet spaces and headsets. A hybrid fallback uses one room audio connection, individual laptops without audio, and a remote advocate. Edit this diagram on Excalidraw

Another arrangement is one room audio connection with individual laptops joining without audio for chat and questions. Zoom warns that muting microphones alone leaves speakers active. Use a shared live agenda and give someone responsibility for bringing remote contributions into the discussion. Breakouts still need physical space and planning.

Separately: record everyone

For me, the bigger opportunity is capturing everyone’s contributions so we can work with them afterwards. A requirements discussion could become material for an agent to extract constraints, find contradictions, ask follow-up questions and build something. That’s much more useful than a summary nobody reads.

The capture needs to include the room, remote participants and breakouts, with participants agreeing to it. Two requirements are easy to confuse: identifying who spoke and saving separate audio tracks. A transcript can name people whose voices still share one recording.

Two people speak through one microphone. Separating their voices in a transcript is different from identifying their names and confirming action owners. Edit this diagram on Excalidraw

The capture problem already has solutions

These are documented product features; their accuracy needs testing in the actual room:

Setup What it already offers What needs configuring
Microsoft Teams speaker recognition Names individual speakers sharing room audio. Microsoft also documents a laptop with a USB speakerphone option. Voice enrolment, meeting invitations and administrator settings. Teams Rooms needs a Pro licence; the laptop host needs Teams Premium or Copilot. Check the calendar requirements.
Zoom Rooms smart name tags for voice Attributes room speech to individuals in captions, transcripts and summaries. Supported room setup, administrator enablement, and voice enrolment/invitations for automatic names. It doesn’t work when the room joins another platform’s meeting.
Ordinary Zoom desktop recording Saves a separate audio file per connected participant, including on the free plan. Enable the separate-audio-files option. Several humans sharing one room connection still share that audio stream.
Otter Distinguishes speakers after processing; naming them helps it recognise their voices in future. Review and correct the names. It also supports audio export, so the transcript needn’t be the only surviving record.

Turn on actual recording as well as transcription: Teams can retain meeting audio, and Zoom cloud recording offers mixed or per-participant audio files. Check retention settings. Breakouts require extra care: Zoom cloud recording captures only the main room; recording each breakout needs a participant recording locally in it.

So this doesn’t require waiting for future technology. The work is choosing and configuring a setup, then checking that it captured everyone correctly. Better attribution still doesn’t make room participants notice a remote colleague trying to interrupt.

Keep the recording and approve the outputs

As transcription models improve, we can revisit the same audio and potentially recover words or speaker distinctions an earlier model missed. A summary alone loses that opportunity: it has already selected what mattered and discarded the rest. Better models still can’t guarantee recovery of speech the microphones never captured clearly.

There is still a reason to synthesise early. People need a concise account of decisions, owners and open questions to review and approve while the meeting is fresh. I’d keep both: the source recording for reprocessing, and the approved outputs for acting on. Better transcription can inform a correction; it shouldn’t silently change what was agreed.

Does Granola solve this?

Granola is more than a shared notepad. It transcribes microphone and computer audio without a meeting bot, combines the transcript with your notes to produce AI-enhanced notes, and makes meeting material available through MCP.

But its desktop speaker tags cannot distinguish people sharing a meeting-room device. Its mobile app can distinguish speakers face-to-face, yet Granola doesn’t save the audio. There is no recording to replay or re-transcribe.

That makes Granola useful for notes, but a poor fit for my requirement to preserve the source. I’d first test the recording and speaker-recognition features in the meeting platform we already use, then check the names, decisions and owners with the people who were there.