GUIDE

How is simultaneous interpretation cost calculated? The language-hour model as an alternative to headset rental

In three sentences

  • In classic simultaneous interpretation cost grows with languages × interpreters × days, plus booth, transmitter, receiver/headset rental and a technical team per hall — audience size enters the invoice through the number of headsets.
  • In the AI-based model the unit is the language-hour: broadcasting one hour of speech into two target languages uses two language-hours; audience size scales the broadcast, not the interpretation cost.
  • For a sound budget look at three numbers: speaking hours, number of target languages, number of halls; when phones and QR replace headset distribution the logistics line almost disappears.

The classic structure: interpreter days, booth, headsets

A booth setup has three main lines. First, interpreter fees: two interpreters per language at a daily rate; a half day usually counts as a full day. Second, technical rental: sound-proof booth, transmitter, language channels and a receiver/headset per attendee; the number of headsets grows one-to-one with attendees and staff are needed to hand them out and collect them. Third, repetition per hall: at a seven-hall fair the booth and team are set up seven times. Most of the cost in this structure is logistics, not the interpretation itself.

The language-hour model: how it counts

In the AI-based setup the unit is the language-hour. One hour of speech broadcast into one target language uses one language-hour; into two languages, two; a two-hour session into two languages, four. VerbaStage plans count exactly this way: source minutes × number of target languages; attendee capacity is a separate dimension of the plan and, within that capacity, audience size does not consume extra minutes. Breaks and interruptions do not count as speaking time.

Why audience size does not raise the cost

One broadcast is produced per target language and reaches everyone in that language room at the same time. Fifty or five hundred listeners hear the same broadcast; that is a broadcast/infrastructure question, not an interpretation invoice. In the classic setup every listener means a receiver and a headset. When attendees use their own phones and earphones this line disappears; planning a small number of loaner earphones for those who need them is enough.

Example scenarios

A one-day seminar, 4 hours of speech, targets English and German: 8 language-hours. A fair with three halls, 6 hours per hall, one target language each: 18 language-hours; in a booth setup that would mean three booths, six interpreters and three equipment sets. A two-day congress, 6 hours a day, three target languages: 36 language-hours. These numbers are the main budget input; the plan is chosen by source hours and capacity.

What deserves extra budget

A good microphone and a clean mixer output on stage; stable internet and a backup connection; a rehearsal on real devices before doors open. These determine interpretation quality and are usually already in the event technical plan. For high-stakes sessions (legal, medical, negotiation) budget human interpreters or a mixed setup; VerbaStage states this boundary openly.

Where do I see prices?

VerbaStage plans and current tax-inclusive prices are published on the plans page. This guide is not a price list; it gives the calculation logic so you can compare offers in the same unit.

Frequently asked

Does a half-hour session count as a full hour?
No; usage follows speaking time. Breaks and interruptions are not counted.
Are headsets never needed?
Attendees use their own earphones. Planning a small number of loaner earphones for those without is enough; the attendee help desk manages it.
Are multiple halls very expensive?
Each hall counts its own speaking hours × languages; without booth and equipment repetition the logistics cost does not grow with the number of halls.
Compared with human interpreters?
The classic setup costs languages × interpreters × days + equipment + logistics; the AI setup counts language-hours. Budget human interpreters for high-stakes content; a mixed setup is the balanced answer for most congresses.

Instead of reading, listen: a live demo with your own voice.

Start a short live demo with a few details and microphone permission; see the delay, audio and captions on your own phone.

Last updated: 2026-08-18. This guide describes real product behaviour; figures follow our realism principle.