The Best AI Note Taker for In-Person Meetings Isn't a Meeting Bot
The short answer: if your meetings happen in a room, rule out anything that works by joining the call as a bot, because there is no call for it to join. What you need is a tool that records locally on a device in the room and separates speakers by voice rather than by who is logged into a meeting. That second requirement is the one buyers miss, and it's the one that decides whether your transcript says who said what.
(Disclosure: Teneks is our product. It does in-person capture, which is why we've tested this carefully. Each tool shape below is the right answer for someone.)
Why meeting bots don't apply
Most AI note takers are built around video calls. They join Zoom, Teams or Google Meet as a participant, receive one audio stream per attendee, and label the transcript from the invite list. That design gives them accurate speaker names for free. It also makes them useless for a conversation in a room, a car, or across a desk. There is no meeting to join and no participant list to read.
For in-person, the field narrows to two shapes:
- Recorder apps on a laptop or phone: start, stop, upload, transcribe
- Dedicated devices: a pendant, badge or pocket recorder that does the same with better microphones
Both work. Both then run into the same problem.
The problem nobody tests for: two voices, one microphone
When a single microphone captures several people, someone has to work out which words belong to whom. There are two ways to do it, and the difference is invisible until it isn't:
- By participant list: reliable, free, and completely unavailable in a room
- By voice: the tool clusters the audio by vocal characteristics and assigns each cluster a speaker
We ran a controlled test where two participants shared one laptop while two others joined remotely. The bot-based tool produced no separate speaker for the second person in the room. His words were folded into the person beside him. A tap-to-record app produced three speakers where there were four. Our own capture separated all four.
Our sample is small, but the mechanism holds: if a tool has no way to tell voices apart, a second person at the same microphone gets absorbed into the first, silently. Nothing in the transcript flags it. You get a clean, readable, confidently wrong document.
Devices vs. apps
Dedicated recorders sell on microphone quality, and for a noisy room that's a real advantage. But the microphone is only half the job:
| Factor | Dedicated device | Phone/laptop app |
|---|---|---|
| Microphone quality in a noisy room | Better | Adequate at close range |
| One more thing to charge and carry | Yes | No |
| Speaker separation | Depends on the software behind it | Depends on the software behind it |
| Works for online meetings too | Usually a separate flow | Often the same tool |
Before paying for hardware, check what the software behind it does with speakers. A better microphone feeding a system that labels by participant list still cannot tell two people apart.
What to look for
1. It records locally, not by joining a call. Test it with the internet off, then let it upload.
2. It separates speakers by voice, and says so. Vendor pages describe this as "speaker diarization", "speaker detection" or "speaker labels". If the wording is about "meeting participants", it's probably reading the invite.
3. It recognises the same voice next time. Otherwise every meeting starts at "Speaker 1", and you rename people forever.
4. The labels survive the export. Some tools ship speaker names in one export format and none in another. Same recording, two very different files. Export what your workflow actually consumes and check.
5. It handles your language, specifically the names. English-first systems transcribe non-English conversation reasonably and then mangle every proper noun in it.
6. It starts before the conversation does. In-person meetings begin the moment you sit down. The useful context and the consent question are both in the first two minutes.
Test it in one meeting
Take your next in-room meeting with at least three people, two of them on the same device.
- Start recording before the greeting.
- Have each person say their name once, early on. It gives you something to check against.
- Halfway through, have the two on the shared microphone speak in quick succession.
- Export and read the result: is everyone there? Did anyone get merged? Are the names right?
Ten minutes of reading tells you more than any feature comparison, including this one.
Where Teneks fits
Teneks records locally on macOS and Windows, separates speakers acoustically rather than from an invite list, learns a voice so the same person is named in later meetings, and works in Estonian, Finnish, Latvian, Lithuanian, Polish and 100+ other languages, including several inside one conversation. Then it produces a summary, decisions and next steps rather than a raw transcript. See Teneks Meetings.
The gap on our side: there's no native iPhone or Apple Watch app today, so phone-recorded conversations come in through upload rather than a dedicated mobile client.
FAQ
What is the best AI note taker for in-person meetings?
One that records locally on a device in the room and separates speakers by voice rather than from a meeting invite. Bot-based note takers cannot attend an in-person meeting at all.
Can Otter or Fireflies record in-person meetings?
Recorder apps can capture in-room audio, but tools designed around joining video calls lose their main advantage in a room: the participant list they use for speaker names doesn't exist.
Do I need a dedicated AI note taker device?
Only for difficult acoustics. The microphone helps in a noisy room, but it does nothing for speaker separation, which is decided by the software behind the device.
How does an AI note taker know who is speaking without a meeting invite?
By clustering the audio into distinct voices and assigning each a speaker, optionally matching those voices to people it has heard before.
Try it on a call of your own
Everything above is easy to claim and easy to check. Drop in a real recording — your language, your speakers, your background noise — and read the transcript and summary of the first 30 seconds. Free, no account.
Test your recordingWritten by Arne Niitsoo
Arne is the founder of Teneks, a call intelligence platform built in Tallinn, Estonia. His work lives inside real conversations: how speech becomes text, how voices get told apart, and why the languages of the Nordics and Baltics break most tools. He writes from what he hears in real calls, not from theory.