How user researchers use an AI note taker for interviews and synthesis
For a researcher, the value of an AI note taker is not that it saves you typing. It is that twelve conversations come back in the same shape, with timestamps, so the pattern across them is something you can check instead of something you remember. This page covers how to set that up: consent first, then capture, then a fixed set of fields per session, then synthesis, then a plain list of what the tool will not do for you.
Updated September 2026
The research week this page is about
A study rarely spreads itself out politely. You recruit eight to twelve participants and most of them land in the same three days. Sessions run 45 to 60 minutes. On a heavy day you run four, twenty minutes apart, and you still owe a stakeholder readout by Friday.
The usual setup is a moderator guide on one screen, the participant on the other, and you typing while they talk. Sometimes a colleague takes notes so you can moderate properly. Often nobody is free and you do both jobs badly at once.
The cost shows up in the same places every time. Session one has rich notes and session eight has four lines, because by then you are tired and you have heard it before. The best moment usually arrives around minute 38, when the participant relaxes, which is exactly when your typing falls behind. Writing the readout, you need one exact sentence and you write down what you think they said. A stakeholder asks how many participants hit the problem and the honest answer is most of them, I think.
The bigger problem only appears at synthesis: your notes are not in the same shape. Participant 2 has a paragraph about pricing because pricing was fresh in your mind that morning. Participant 9 has nothing about pricing, and you cannot tell whether they never mentioned it or you never wrote it down. Recording every session fixes that. The transcript is not the deliverable. The consistency is.
Noter AI records, transcribes and summarizes your meetings on iPhone, iPad & Android, in 60+ languages.
Consent and participant data come first
Ask on the recording, at the start, in plain words, and let their yes be the first thing on the file. Then say what happens to it afterwards, specifically, because a participant who agrees without understanding has not really agreed.
Research consent needs more care than a standup recording for two reasons. You are paying most participants, which quietly makes it harder for them to say no. And what they tell you is personal: their job, their money, their health, their household. If your work sits under an ethics board or an IRB, the recording, the transcript and the third-party processing belong in the protocol you submitted, not in a decision you make on the day.
Where the audio goes matters as much as whether you asked. Noter AI processes recordings in the cloud, encrypted in transit and at rest, and you can delete a note and its recording at any time. If you recruit in the EU, read is AI meeting transcription GDPR compliant before the first session. Two more worth reading once: does an AI note taker train AI on your recordings and how to delete a meeting recording permanently.
- A script that works: I record these sessions so I can quote you accurately instead of paraphrasing you. A third-party app transcribes it, I am the only person who reads it, and I delete it when the study closes. Are you comfortable with that?
- Say the incentive does not depend on it. You still get the voucher if you say no. Say it out loud. It changes how honestly people answer.
- Offer a pause. They can stop the recording at any point, including halfway through an answer. Some of the best material arrives right after someone tests whether you meant that.
- Pseudonymise inside the transcript. The transcript is editable, so you can replace a real name with P4 and re-run the summary so the summary uses the code too. Do it the same day.
- Set the deletion date when you recruit, not when you finish. Deleted when the readout ships is a rule. I will clean it up sometime is not.
- In group and dyad sessions, ask everyone, not just the person who booked the slot.
Where research sessions happen, and what captures each
Research capture is not one situation. In a normal study you hit three or four of these, sometimes in the same week.
| Situation | How it happens | What captures it |
|---|---|---|
| Remote moderated session | Zoom, Google Meet, Microsoft Teams or Webex link in the invite | Send the bot. Connect Google Calendar or Outlook and it joins from the invite, so you are not pasting links between back-to-back sessions. |
| Lab or meeting-room session | You and one participant across a table, sometimes an observer | Phone on the table, one tap from the Control Center widget. Recording keeps running with the screen locked. |
| Contextual inquiry or site visit | A shop floor, a clinic, a warehouse, somebody's desk | Same one-tap recording. Put the phone flat and close to the participant, and expect background noise in the transcript. |
| Home visit or diary debrief | Kitchen table, kids in the background, no laptop open | Phone recording. Nothing to set up beyond the consent line. |
| A voice note the participant sends afterwards | WhatsApp voice message, or a memo from their own phone | Share it into the app from the iOS share sheet. It becomes a note like any other, transcribed and searchable. |
| A session a colleague recorded | An audio file, M4A, MP3, WAV, AAC or similar, in Files, iCloud Drive or a shared folder | Import the audio file directly. Same transcript, speaker labels and summary as a live recording. |
- What this does not cover: if your platform hands you a video file of the session, the app will not take it. Audio only. More on that in the limits section.
The part most tools get wrong: participants who do not speak your language
Recruit people in the language they think in and your data gets better. People describe frustration, money and embarrassment far more precisely in their first language. This is also where most transcription tools fall over, and the failure is quiet enough that you may not notice until synthesis.
The common design is that you choose a language before you press record. Fine for a monolingual session, wrong for almost every real one. A participant in Riyadh, Mumbai, Jakarta or Barcelona speaks their own language and drops English words into the middle of the sentence: checkout, coupon code, free trial, refund. A tool locked to one language either writes those words out phonetically in the wrong script, so they spell nothing, or drops them.
That is a synthesis problem, not a cosmetic one. If coupon code was transliterated into letters that spell nothing, searching your twelve transcripts for coupon returns zero hits, and the fact that six participants tripped on the same thing never surfaces. You end up reporting some participants mentioned discounts when you could have reported six of twelve, with a timestamp each.
Noter AI tags language word by word instead of committing the whole recording to one language. A sentence that runs Arabic, then English, then back to Arabic arrives in one transcript with each stretch in its own script. You never pick a language before recording. Transcription covers 60+ languages with detection always on, and a finished note can be translated into 18 languages, so a study run in Turkish produces a readout your English-reading stakeholders can use. The wider comparison lives on the best app for transcribing multilingual meetings.
One practical note. Translate the summary and the quotes you plan to show, and keep the original transcript next to them. When someone questions a quote, you want to point at the participant's own words, not only at a translation of them.
Decide what comes out of every session before you run the first one
The highest-value setting is the output shape. Fix it before session one, use the same one for all twelve, and synthesis stops being an archaeology project.
Noter AI reshapes the same recording into several formats: key bullet points, minutes, to-do list, formal report, email, or a Custom instruction where you describe the structure you want. For research, use Custom, because the fields you need are not the fields a sales meeting needs. The Custom instruction is also where you fix the house voice, worth writing once if your readouts have a voice your team recognises.
Here is a worked example for a checkout onboarding study. Write the Custom style once and every session comes back in this shape:
- Participant context. Role, how long they have used the product, how they found it. Two lines maximum.
- What they were trying to do, in their own words rather than in your feature names.
- What they actually did, in order, including the wrong turns.
- Where they stalled, with the moment it happened so you can go back to the audio.
- Their workaround, if they had one. This is usually where the design insight is.
- Two verbatim quotes, picked for how clearly they say the thing, not for how quotable they sound.
- What they asked me. The questions the participant put back to you. An underrated field.
- Open threads, anything you meant to follow up on and ran out of time for.
- Action items, which the app pulls out on its own: send the voucher, share the prototype, book the follow-up.
Verbatim quotes with timestamps, which is what a readout runs on
A readout is only as strong as the quotes in it. A paraphrase invites a stakeholder to argue with your interpretation. The participant's own sentence does not.
Every transcript segment carries a timestamp, and tapping a line plays the audio from that exact moment. That gives you two things: you can check a quote in three seconds instead of scrubbing a 52-minute recording, and you can put the timestamp beside the quote so anyone who doubts it can go and hear it. That habit changes how research lands in a room. It turns users found this confusing into P4 at 00:31:12 and P7 at 00:19:40.
Speaker labels separate you from the participant, which matters more than it sounds. Half of what you say in a session is a leading question, and seeing your own prompt directly above their answer is a useful check on whether you fed them the insight. Same mechanism described in can AI tell who said what in a meeting.
Fix the transcript before you synthesise. Product names, internal jargon and the participant's own company name are the words a speech model is most likely to get wrong, and they are also the words you will search for later. The transcript is editable, and after you correct it you can re-run the summary so the summary picks up the fix instead of repeating the error. Five minutes per session, done the same day, is the highest-return five minutes in this whole workflow.
Finding the pattern across twelve sessions, without pretending it is coding
Put the study in its own folder and every note becomes searchable by title and AI notes, with a word search inside any transcript you open. That alone answers the did anyone else say this question that used to cost an afternoon.
AI chat goes further: you can ask one question across all your notes at once instead of opening them one by one. Which participants mentioned price before I asked about price? Who abandoned at the address step? List every time someone said they would ask a colleague for help. You get back a set of pointers. More on how that works in asking questions across all your meeting notes.
Now the discipline part, because this is where AI-assisted research goes wrong. Treat everything the chat returns as a search result, not a finding. It tells you where to look. It does not tell you what is true. Before anything from a chat answer reaches a slide, open the note, listen at the timestamp, and read the sentences either side. Half the time the context changes what the line means.
Keep observation and interpretation in separate columns and do not let the tool blur them. The observation column holds what happened and what was said, with a participant code and a timestamp, and a machine can help you fill it. The interpretation column holds what you think it means, carries your initials, and is yours alone. When a stakeholder pushes back, you want to point at the exact line where the two columns meet.
Be honest about the summaries too. A summary compresses one session. Eleven summaries side by side are not a thematic analysis. They are a faster route to one, which is worth paying for, but the analysis is still your job.
A realistic study week, step by step
- 1Monday, before anything. Create a folder for the study. Write your Custom output style once, with the fields above. Confirm your consent script and your deletion date.
- 2Monday. Record the stakeholder intake call the same way you will record participants. On Friday you will want to check what they actually asked for against what you remember.
- 3Tuesday and Wednesday, sessions. Remote sessions: the bot joins from your calendar. In person: one tap, screen locked, phone flat on the table. Ask for consent on the recording every time, even with a participant you met last month.
- 4Five minutes after each session. Fix names and product terms, re-run the summary, swap real names for participant codes, and flag the two quotes you would use if this were the only session you ran.
- 5Wednesday evening. Import anything that arrived sideways: a colleague's audio file, a participant's voice note, a session recorded on another device. Everything belongs in the same folder before you look for patterns.
- 6Thursday morning, first pass. Ask the same three or four questions across the whole folder. Write down the pointers, not the conclusions.
- 7Thursday afternoon, verification. Open every pointer at its timestamp and listen. Keep the ones that survive. Note the ones that turned out to be something else, because those are often more interesting.
- 8Friday, write the readout. Observation column with participant codes and timestamps, interpretation column with your name on it. Paste quotes exactly as spoken. Translate the summary and quotes if your audience reads another language.
- 9Friday, ship it. Export to PDF for the record, Word if a colleague needs to edit, or copy formatted text straight into the document your team already uses.
- 10When the study closes. Delete the recordings on the schedule you promised participants.
What Noter AI will not do for a researcher
This is the section that decides whether the tool fits your study. None of it is softened.
No research repository integration. There is no Dovetail, Condens, EnjoyHQ, Notion, Airtable, Miro or Zapier connection. None at all. You export to PDF, Word or plain text, or copy rich text, and paste. If your process assumes transcripts land in a repository automatically, that is a manual step every session, and you should count the cost before committing.
No automatic thematic coding. No codebook, no tags, no generated themes, no affinity board, no count of how many participants hit a code. You get summaries in a shape you chose, plus chat across every note. That gets you to the pattern faster. It does not do the coding for you.
Cloud processing only. No on-device or offline transcription. Audio is uploaded, transcribed in the cloud, encrypted in transit and at rest, and deletable by you at any time. If your ethics approval or your client's contract forbids sending participant audio to a third-party processor, this is a hard stop, and you need to know it before you recruit rather than after twelve sessions.
No video, in any form. MP4 and MOV files are not accepted. If your platform gives you a video recording of the session, the app cannot take it. Audio only: M4A, MP3, WAV, AAC and the other formats iOS handles natively. There is no URL or link import either.
No screen recording. For a usability test where the click path is the data, this covers the talk track and nothing else. You still need a separate screen capture tool, and you will still line the two up by hand.
No live captions during the session. Transcription happens after you stop recording, not while the participant is talking.
No PDF import. Discussion guides, screener responses and prior reports stay wherever they already are.
No published accuracy number, and field conditions are hard. A quiet remote session transcribes well. A shop floor, a strong accent on a bad connection, or two people talking over each other will need a read-through before you quote anything. Check every quote against the audio before it goes on a slide.
What it costs when you are one researcher, or three
Most research teams are small, which is exactly the shape per-seat pricing punishes. One researcher, or two researchers and a designer who runs the occasional session, or an academic with two RAs. You want everyone who runs a session to be able to record it, and per-seat billing turns that into a budget conversation with someone who does not care about your study.
Noter AI is $9.99 a month or $49.99 a year, per person rather than per seat, so three researchers is three subscriptions at about $30 a month, or about $12.50 a month between them on the annual plan. There is a free trial, so test it on a real session rather than a rehearsal. A rehearsal never has the background noise, the accent, or the mid-sentence language switch that actually matters.
For comparison, using published prices: Otter.ai Pro $16.99/month, Notta Pro $13.99/month, Granola $14/user/month on the Business tier, tl;dv Pro $18/user/month, Fireflies $10 to $19 per seat per month, and Fathom free with Premium at $20/month, Team at $19/user/month and Business at $34/user/month.
The number that decides it is usually not the monthly rate, it is who counts as a user. Research is done by a rotating cast: a researcher who runs sessions every week, a designer who moderates two during a discovery sprint, a PM who takes the notes when nobody else is free. On a per-user plan you either buy licences for people who moderate three times a year or you lock them out and lose the recording. Here each of them records from their own phone on their own subscription, so the occasional moderator costs one month, not one seat.
Available for iPhone and iPad and for Android, and notes sync across your devices, so a session recorded on a phone in the field is on your iPad when you sit down to synthesise. For a single session end to end, see how to record and transcribe an interview. For the fields themselves, free and with nothing to install, take the user interview notes template.
Frequently asked questions
What is the best AI note taker for user interviews?
The one that captures every situation your study hits and returns the same fields every time. For most researchers that means recording in person as well as on a call, accepting an audio file a colleague sends, timestamping every line so you can verify a quote, and coping with a participant who does not speak your language. Noter AI does all four at $9.99 a month flat. It does not do thematic coding or repository integration, so check those against your process first.
Can it transcribe an interview where the participant switches between two languages?
Yes, and this is the sharpest reason to use it for international research. Language is tagged word by word rather than fixed for the whole recording, so a participant speaking Hindi who drops English product words into the sentence is quoted accurately on both. You never choose a language before recording. Tools that lock to one language at record time write the second language out phonetically or drop it, which quietly ruins any search you run across your transcripts later.
Do I have to tell participants I am recording?
Yes, always, and get the yes on the recording itself. Beyond the legal question, which varies by country, research consent has its own standard: participants are usually paid, and they are telling you personal things. Say who will hear it, that a third-party app transcribes it, and when you will delete it. Make refusing easy and confirm the incentive does not depend on it.
Can it tag themes or push transcripts into Dovetail?
No to both. There is no integration with Dovetail, Condens, Notion, Airtable or any other repository, and no automatic thematic coding, tags or codebook. You export to PDF, Word or plain text, or copy rich text, and paste into whatever you use. What you do get is chat across every note at once, which is a fast way to find where a theme lives before you code it yourself.
Can I get verbatim quotes with timestamps for a readout?
Yes. Every transcript segment carries a timestamp and tapping a line plays the audio from that exact moment, so checking a quote takes seconds. Speaker labels separate your questions from the participant's answers, which is a useful check on whether you led them. Put the participant code and the timestamp next to each quote so anyone who doubts it can go and hear it.
Can I import a session a colleague already recorded?
Yes, as long as it is audio. Import M4A, MP3, WAV, AAC or any other audio format iOS recognises from the Files app or iCloud Drive, or share a file in from another app through the iOS share sheet, including a participant's WhatsApp voice note. Video files are not accepted in any format, so a session saved as an MP4 cannot be brought in.
Is it usable for a usability test?
For the talk track, yes. For the click path, no. There is no screen recording, so you still need a separate tool to capture what the participant does on screen, and you will line the two up by hand. Where it helps is the spoken layer: what they expected, where they hesitated, and the exact sentence at the moment they gave up.
How much is it for a small research team?
$9.99 a month or $49.99 a year, flat, not per seat. The comparison worth running is per study rather than per month: a round of eight interviews on the annual plan costs roughly one participant incentive. Nothing inside the plan is metered either, so a heavy fieldwork fortnight does not run you out of summaries. There is a free trial, so run it on a real session with real background noise rather than a rehearsal.
Let Noter AI take your meeting notes
Record, transcribe, and summarize meetings on iPhone, iPad & Android, or send a bot to Zoom, Teams, Meet, or Webex. In 60+ languages.
On your computer? Scan with your phone to get the app.