Market Research Transcription Cost: Focus Groups, IDIs, and Multi-Speaker Sessions
Date Published

Quick answer: Market research transcription is usually easiest to budget by audio minute or audio hour, then adjust for the work that makes a file harder: more speakers, heavy crosstalk, poor audio, full-verbatim detail, timestamping, specialized terminology, bilingual transcription/translation, anonymization, and rush turnaround. A 60-minute one-to-one IDI is not operationally equivalent to a 60-minute eight-person focus group. Ask vendors to price the actual session type, not just duration.
Market research teams often discover transcription cost after fieldwork is locked. By then, a project may include 24 IDIs, six 90-minute groups, three markets, rush topline analysis, and a client that wants every quote tied back to video. The best time to design the transcript is when the discussion guide is being finalized.
VerbalScripts provides focus group and interview transcription and solutions for market researchers. Request a current project quote using the formula below.
Base formula: total recorded minutes × applicable rate
Calculate each methodology separately.
Example project:
• 20 IDIs × 60 minutes = 1,200 minutes;
• 4 focus groups × 90 minutes = 360 minutes;
• 6 stakeholder interviews × 45 minutes = 270 minutes;
• total = 1,830 recorded minutes.
Then ask whether the vendor applies different rates for one-to-one versus multi-speaker files. Some services use one rate but add options; others quote the group work separately.
Why focus groups cost more to produce well
Focus groups require ongoing speaker identification. The transcriptionist must distinguish comments such as:
P03: I’d buy it if the price stayed under fifty.
P07: I actually think fifty makes it look cheap.
If both comments become Speaker:, or are assigned to the wrong participant, the transcript loses analytical value.
Cost factors include:
• 6-10 voices with similar acoustic profiles;
• overlap and laughter;
• people talking from different distances;
• moderator interruptions;
• participant names/codes;
• stimulus references;
• side conversations.
For operational guidance, see Healthcare Focus Group Transcription; many of the same speaker-control principles apply outside healthcare.
IDIs are simpler - until they are not
A well-recorded 1:1 depth interview is a straightforward two-speaker task. Costs can rise when it contains:
• strong background noise;
• very soft telephone audio;
• specialized B2B terminology;
• product/model names;
• accents unfamiliar to a generalist;
• two interpreters or observers joining;
• exact timestamp requirements;
• bilingual content.
Do not assume “IDI” means “easy.” Send a representative sample if audio conditions vary.
Eight variables to put in the quote request
1. Number and duration of files
Provide actual or expected audio minutes, not just “30 interviews.”
2. Speaker count
State the maximum and average speakers per session.
3. Speaker-identification requirement
Does P01/P02 need to remain stable across the entire session? Will the vendor receive a roster or multi-track recording?
4. Verbatim level
Full verbatim preserves fillers, false starts, and more nonverbal activity. Clean verbatim reduces routine disfluency while preserving meaning.
5. Timestamps
Fixed interval, every speaker, every moderator question, or only uncertain passages? More granular timestamps add work but can save analyst time.
6. Anonymization
If participant names, brands, client names, or PII must be replaced, define the rule and whether a restricted master is also delivered.
7. Language
Separate transcription and translation in the quote. A Spanish transcript plus an English translation is a two-stage deliverable. See Spanish-English Research Interview Transcription.
8. Turnaround
A project arriving Friday night for Monday morning analysis is a different capacity request from a rolling five-day schedule.
Calculate the cost of analyst time too
A cheap transcript that requires the research team to repair every speaker label can be expensive in practice. Compare:
Total project cost = transcription fee + analyst cleanup time + quote verification time + rework
For every pilot, record:
• minutes spent correcting each hour of transcript;
• speaker-attribution errors;
• missed brand/product terms;
• unusable overlap;
• time to find a quote in source audio;
• client-requested corrections.
The better vendor is the one that reduces the total work needed to turn fieldwork into findings.
Ethics and confidentiality belong in the budget conversation
AAPOR’s standards and the ICC/ESOMAR Code emphasize research integrity, transparency, participant responsibility, privacy, and professional conduct. A transcription vendor becomes part of the research data chain.
Ask who has access, how files are transferred, whether subcontractors are used, and when recordings/transcripts are deleted. For sensitive categories, client NDAs and participant commitments may be stricter than a generic platform’s default terms.
How to lower cost without harming the data
1. Improve recording quality. Separate mics or tracks can reduce correction effort.
2. Use participant codes from the start. Do not make a vendor infer a speaker map after fieldwork.
3. Provide a glossary. Brands, competitor names, study stimuli, and technical language should be supplied once.
4. Select timestamps strategically. If analysts only need to retrieve quotes, speaker-turn or periodic timestamps may be enough.
5. Use rolling delivery. Avoid creating artificial rush pressure at the end.
6. Translate only what the methodology requires. If bilingual analysts code source language, selected-quote translation can be more efficient.
7. Pilot one file. Validate formatting before 40 files use the wrong convention.
A practical budgeting table
1:1 IDI — Main cost driver: Duration + terminology | What to specify: 2 speakers, glossary, verbatim level
Dyad — Main cost driver: Overlap + 3 voices | What to specify: stable labels
Focus group — Main cost driver: Speaker attribution + crosstalk | What to specify: participant roster/track layout
Online community video — Main cost driver: variable audio + many clips | What to specify: file naming and timestamps
B2B expert interview — Main cost driver: specialist terms | What to specify: industry glossary
Multilingual group — Main cost driver: transcription + translation | What to specify: language-by-language deliverables
Use a pilot to estimate the real rate for multi-speaker work
Before committing a large study, select a representative 10- to 20-minute sample that includes the conditions most likely to raise transcription effort: overlapping speakers, brand names, moderator probes, remote participants, or a mix of accents. Ask the vendor to return the sample using the exact speaker labels, timestamp convention, and anonymization rules planned for the full project.
Then measure more than price. Record how much analyst cleanup is still required, whether respondent quotes can be traced back to the audio quickly, whether speaker identities remain consistent, and how many terminology corrections are needed. A slightly higher transcription rate can be cheaper overall if it removes hours of project-manager and analyst rework across dozens of sessions.
Frequently asked questions
How much does focus group transcription cost?
Pricing varies by vendor and project. Focus groups often cost more than 1:1 interviews at the same duration because speaker identification and overlap require more work. Request a quote using real speaker count and audio samples.
Is transcription billed by audio minute or working hour?
Many services bill by recorded minute/hour. Confirm the unit. A 90-minute focus group remains 90 audio minutes even if it takes several working hours to transcribe and review.
Does timestamping cost extra?
It can. The price depends on how frequently timestamps are required and whether speaker labels must also be verified.
Is AI enough for market research?
AI can be useful for quick drafts, but multi-speaker groups, accents, brands, and crosstalk can degrade speaker attribution and word accuracy. Human review is valuable when the transcript drives client findings and quotations.
Can VerbalScripts handle recurring research agency volume?
Yes. Provide expected weekly/monthly audio, session types, formatting, security terms, and delivery SLA when requesting a quote.
Price the transcript you actually need
A useful market-research quote is not “$X per minute.” It is “$X for a transcript with the speaker accuracy, timestamps, privacy controls, language handling, and turnaround our analysts need.” Send VerbalScripts your fieldwork plan before launch to scope the real cost.
Authoritative references
• AAPOR Standards and Ethics (Code revised in 2026)