We Recorded One Meeting With Three AI Note Takers. Here's What Each Missed.
We recorded the same 57-minute meeting with three AI note takers and compared the transcripts. All three captured useful discussion, but they differed on the opening minutes, speaker labels and several Estonian and Finnish names.
Teneks is our product and was one of the tools in the test. This article reports what happened in that recording, with the measurement limits explained below.
The setup
One real meeting. Four participants across three locations, two of them sharing a single laptop in the same room. Estonian and English inside the same conversation, plus Finnish and Estonian company names. Google Meet, 57 minutes.
Three tools recording simultaneously:
- A meeting bot that joins the call as a participant
- A tap-to-record app running on one participant's machine
- Our own desktop capture (Teneks)
Because each tool starts its own clock at a different moment, the first job was aligning them. We anchored on the first sentence all three contain and recovered the offsets by maximising content-word overlap. Everything below is measured from that shared anchor forward, so no tool is judged on a window it never had.
How we compared the recordings
We could measure objectively:
- Coverage: how much of the meeting exists in each transcript
- Name accuracy: verifiable against the audio
- Turnaround: timestamps on delivery
- Speaker agreement: how often two tools say the same person was talking
We could not measure accuracy of speaker labels, because there is no human-annotated ground truth for this recording. When we report "94.7% agreement", that means agreement with our transcript, not correctness. That limits the conclusions we can draw from it. We therefore focus on the capture timing and errors we could check against the audio.
Finding 1: two of the three missed the first three minutes
| Tool | Captured | Missing |
|---|---|---|
| Desktop capture | 0:00 – 57:04 | nothing |
| Meeting bot | 2:54 – 57:02 | first 2 min 54 s |
| Tap-to-record app | 2:56 – 57:04 | first 2 min 56 s |
Those 174 seconds contained 54 turns from all four participants, including the exchange where one person asks whether it's okay to record and the others agree.
The two causes are different and worth separating. The bot starts when it is admitted to the call, and someone has to let it in. On this recording the host was still granting it permission almost three minutes in. The tap-to-record app starts when a human taps record, which on this call happened at almost the same moment. That one is an operator choice, not a product limitation.
In this meeting, both later starts missed the greeting and the recording-consent exchange. Capture timing is worth checking separately from transcript quality. Our article on recording start times looks more closely at that part of the test.
Finding 2: the person who wasn't on their own microphone
Two participants shared one laptop. In the bot's transcript, one of them does not exist. His words were attributed to the person sitting next to him. The tap-to-record app also produced no separate speaker for him. His few lines went to a third participant.
A participant list alone cannot identify two people sharing one login and microphone. Separating their contributions requires distinguishing the voices in the room. This test showed an attribution problem, but it did not isolate how each tool produced its speaker labels.
One caveat: on this recording the fourth participant spoke 17 of his 20 turns inside the first 2:54, the window the other two tools never recorded. In the stretch all three captured he says about five seconds' worth. That short overlap limits what this meeting tells us about speaker separation. We would need more recordings to assess how the tools handle two people sharing a microphone. If in-person is your use case, test it yourself. It takes one meeting.
Finding 3: errors in names and specialist terms
Every tool transcribed the English discussion competently. Then the conversation touched Estonian and Finnish proper nouns:
| Spoken | Tool A | Tool B |
|---|---|---|
| A Finnish first name | rendered as an English name | correct |
| An Estonian first name | rendered as an unrelated English word | correct |
| An Estonian company name | two unrelated English words | close, but the Finnish spelling |
| A Nordic retail brand | a different English word | correct |
| The word "notetaker" | "paper" | "note paper" |
These errors are easy to overlook when the surrounding text reads naturally. Check names against the audio before using a summary for follow-up. Include local company names and product terms in your own evaluation; our multilingual guide suggests what to look for.
Finding 4: turnaround, in context
Our output (transcript, summary, decisions and next steps) was ready a median of about six minutes after meetings of this length ended, measured across a month of production calls. The following turnaround figures were collected for the original article. They mix vendor claims with a reviewer report and were not measurements from this three-tool test:
| Tool | Published turnaround |
|---|---|
| Fathom | under a minute (marketing claim) |
| Otter | transcript 1.5–3 min, summary 2.5–6 min |
| Fireflies | "typically complete within 5 to 10 minutes" |
| tl;dv | 10–15 min (reviewer-reported) |
| MeetGeek | no figure published |
Decide how quickly your team needs the result. Treat any vendor's number as a typical case, not a guarantee. Ours included.
What to check in your own evaluation
- Ask when the tool starts recording, not how accurate it is. Check capture timing separately from transcription quality: audio that was not recorded cannot appear in the transcript.
- If any of your meetings happen in a room, test that case specifically. Check whether each person has a separate, correct label when they share a microphone.
- If your business runs in a language other than English, test the names. Include company names, product terms and figures as well as the general wording.
- Check which export you'll actually consume. One tool in this test ships a labelled transcript in one format and an unlabelled subtitle file in another, from the identical recording. If your workflow is wired to the wrong one, you lose the speakers without noticing.
How to run this test yourself
You need one real meeting and about an hour.
- Record the same meeting with two or three candidates at the same time. Include the join and greeting phase, because that's part of the test.
- Export from each. Note which export format carries speaker labels.
- Align the clocks. Pick the first sentence all transcripts contain and treat it as t=0 for each.
- Compare from that anchor forward: what's in the head that only one tool has, what happens to each named person, and whether anyone is missing.
- Judge the transcript you'd actually paste into an AI assistant, not the one in the vendor's UI.
Keep the original exports and your comparison notes so you can trace each finding back to the recording.
FAQ
Which AI note taker is most accurate?
This test cannot rank overall accuracy. It found differences in capture timing, speaker attribution and names within one meeting. Use those findings to choose checks for your own recordings.
Do AI meeting bots record the whole meeting?
Not necessarily. A bot only records from the moment it is admitted to the call. In our test that was 2 minutes 54 seconds after the conversation started, which meant the recording-consent exchange was not in its transcript.
Can an AI note taker tell apart two people sharing one microphone?
It needs to distinguish the voices rather than rely only on the participant list. Check the output on a shared-microphone recording; this test alone does not establish how each product performs across rooms.
How long should an AI note taker take to produce notes?
Published figures range from under a minute to about fifteen. Whether that is fast enough depends on your workflow. Treat vendor figures as typical rather than guaranteed.
Try it on a call of your own
Everything above is easy to claim and easy to check. Drop in a real recording (your language, your speakers, your background noise) and read the transcript and summary of the first 30 seconds. Free, no account.
Test your recordingWritten by Arne Niitsoo
Arne is the founder of Teneks, a call intelligence platform built in Tallinn, Estonia. His work lives inside real conversations: how speech becomes text, how voices get told apart, and why the languages of the Nordics and Baltics break most tools. He writes from what he hears in real calls, not from theory.