All articles
Tools & Comparisons|9 min read||By Arne Niitsoo

We Recorded One Meeting With Three AI Note Takers. Here's What Each Missed.

The short version: we recorded the same 57-minute meeting with three AI note takers at once and compared the transcripts line by line. All three produced a usable record of what was discussed. They differed on three things that no feature page mentions: how much of the meeting they captured at all, what happened to the person who wasn't on their own microphone, and what they did to names that aren't English.

(Disclosure: Teneks is our product and one of the three tools. Where we used our own output as the reference, we say so, and it changes how you should read those numbers.)

The setup

One real meeting. Four participants across three locations, two of them sharing a single laptop in the same room. Estonian and English inside the same conversation, plus Finnish and Estonian company names. Google Meet, 57 minutes.

Three tools recording simultaneously:

  • A meeting bot that joins the call as a participant
  • A tap-to-record app running on one participant's machine
  • Our own desktop capture (Teneks)

Because each tool starts its own clock at a different moment, the first job was aligning them. We anchored on the first sentence all three contain and recovered the offsets by maximising content-word overlap. Everything below is measured from that shared anchor forward, so no tool is judged on a window it never had.

What we measured, and what we did not

We could measure objectively:

  • Coverage: how much of the meeting exists in each transcript
  • Name accuracy: verifiable against the audio
  • Turnaround: timestamps on delivery
  • Speaker agreement: how often two tools say the same person was talking

We could not measure accuracy of speaker labels, because there is no human-annotated ground truth for this recording. When we report "94.7% agreement", that means agreement with our transcript, not correctness. On our own blog, with our own tool as the yardstick, that number deserves your scepticism. That's why the findings that matter below don't rest on it.

Finding 1: two of the three missed the first three minutes

ToolCapturedMissing
Desktop capture0:00 – 57:04nothing
Meeting bot2:54 – 57:02first 2 min 54 s
Tap-to-record app2:56 – 57:04first 2 min 56 s

Those 174 seconds contained 54 turns from all four participants, including the exchange where one person asks whether it's okay to record and the others agree.

The two causes are different and worth separating. The bot starts when it is admitted to the call, and someone has to let it in. On this recording the host was still granting it permission almost three minutes in. The tap-to-record app starts when a human taps record, which on this call happened at almost the same moment. That one is an operator choice, not a product limitation.

The consequence is the same either way: the beginning of a meeting is the least-recorded part of it, and the beginning is where greetings, context and consent live. If your compliance policy says the recording must contain the consent question, a tool that starts after the greeting cannot satisfy it. We wrote about that separately in why the opening minutes decide your recording compliance.

Finding 2: the person who wasn't on their own microphone

Two participants shared one laptop. In the bot's transcript, one of them does not exist. His words were attributed to the person sitting next to him. The tap-to-record app also produced no separate speaker for him. His few lines went to a third participant.

The reason is architectural. A meeting bot receives one audio stream per participant and labels by who is logged in, which is why its names are usually right and free. Nobody is logged in for the second person in the room, so there is no label to assign. Separating them means telling the voices apart acoustically. That is a different technique, and it's the same one you need for a conversation in a car, a shop floor, or a phone lying on a table.

One caveat: on this recording the fourth participant spoke 17 of his 20 turns inside the first 2:54, the window the other two tools never recorded. In the stretch all three captured he says about five seconds' worth. So this meeting demonstrates the mechanism clearly, but it doesn't prove how those tools would handle a room where the second person talks constantly. If in-person is your use case, test it yourself. It takes one meeting.

Finding 3: names stop working outside English

Every tool transcribed the English discussion competently. Then the conversation touched Estonian and Finnish proper nouns:

SpokenTool ATool B
A Finnish first namerendered as an English namecorrect
An Estonian first namerendered as an unrelated English wordcorrect
An Estonian company nametwo unrelated English wordsclose, but the Finnish spelling
A Nordic retail branda different English wordcorrect
The word "notetaker""paper""note paper"

This is the failure mode that matters most in the Nordics and Baltics, and it's invisible in an English demo. A transcript that reads perfectly can still put the wrong name against every action item. A memo with the wrong name in it is worse than no memo, because someone will act on it. We covered the underlying reason in why most sales AI fails outside English.

Finding 4: turnaround, in context

Our output (transcript, summary, decisions and next steps) was ready a median of about six minutes after meetings of this length ended, measured across a month of production calls. For context, here is what other vendors say in their own documentation:

ToolPublished turnaround
Fathomunder a minute (marketing claim)
Ottertranscript 1.5–3 min, summary 2.5–6 min
Fireflies"typically complete within 5 to 10 minutes"
tl;dv10–15 min (reviewer-reported)
MeetGeekno figure published

Nobody in this category is slow enough for speed to decide a purchase. Treat any vendor's number as a typical case, not a guarantee. Ours included.

What we'd actually tell a buyer

  1. Ask when the tool starts recording, not how accurate it is. Accuracy differences between the leading tools are small. Coverage differences are total: a sentence nobody recorded is 0% accurate in every tool.
  2. If any of your meetings happen in a room, test that case specifically. Invite-based speaker labels are excellent right up until two people share a microphone, then they fail silently.
  3. If your business runs in a language other than English, test the names. Not the grammar, the names. That's where the damage concentrates.
  4. Check which export you'll actually consume. One tool in this test ships a labelled transcript in one format and an unlabelled subtitle file in another, from the identical recording. If your workflow is wired to the wrong one, you lose the speakers without noticing.

How to run this test yourself

You need one real meeting and about an hour.

  1. Record the same meeting with two or three candidates at the same time. Include the join and greeting phase, because that's part of the test.
  2. Export from each. Note which export format carries speaker labels.
  3. Align the clocks. Pick the first sentence all transcripts contain and treat it as t=0 for each.
  4. Compare from that anchor forward: what's in the head that only one tool has, what happens to each named person, and whether anyone is missing.
  5. Judge the transcript you'd actually paste into an AI assistant, not the one in the vendor's UI.

Steps 1–3 are the ones people skip, and they're the ones that produce the findings.

FAQ

Which AI note taker is most accurate?

On English speech in a video call, the leading tools are close enough that accuracy is rarely the deciding factor. The differences that showed up in our test were coverage (what got recorded at all), speaker separation when two people share a microphone, and handling of non-English names.

Do AI meeting bots record the whole meeting?

Not necessarily. A bot only records from the moment it is admitted to the call. In our test that was 2 minutes 54 seconds after the conversation started, which meant the recording-consent exchange was not in its transcript.

Can an AI note taker tell apart two people sharing one microphone?

Only if it separates speakers by voice. Tools that label speakers from the meeting's participant list have no label to assign to a second person in the room, so that person's words are attributed to whoever is logged in.

How long should an AI note taker take to produce notes?

Published figures range from under a minute to about fifteen. Anything in that band is fast enough for normal use. Treat vendor figures as typical rather than guaranteed.

Try it on a call of your own

Everything above is easy to claim and easy to check. Drop in a real recording — your language, your speakers, your background noise — and read the transcript and summary of the first 30 seconds. Free, no account.

Test your recording

Written by Arne Niitsoo

Arne is the founder of Teneks, a call intelligence platform built in Tallinn, Estonia. His work lives inside real conversations: how speech becomes text, how voices get told apart, and why the languages of the Nordics and Baltics break most tools. He writes from what he hears in real calls, not from theory.