AI dubbing with dialogue editing and mix — Warsaw studio
AI dubbing services in a Warsaw studio: script adaptation, sync, dialogue editing and mix — with a documented origin for every voice we use.
AI dubbing usually comes down to a single click and a file that arrives a few minutes later. If you have material that has to go out in several languages — a training course, a corporate film, a documentary, a video channel — you have probably tried a tool like that already, and something in the result grated, even if it was hard to say what. At our sound post-production studio in Warsaw (Żoliborz) we start work exactly where the generator stops.
The model delivers a first pass — dubbing starts after that. We list all sixteen steps that separate the two, before we quote a single one of them. A raw generation is material, not a product: it has a timbre and a script, but no sync, no breaths, no acoustics of the room the character is standing in, and no place of its own in a mix alongside music and effects.
The work is run by Olga Pasternak and Yohan Boisgontier — the same people who edit dialogue and run ADR sessions on films and series. We treat AI dubbing as a language version made in a studio: with the same craft we bring to voice-over and voice recording , only with the voice material sourced differently.
What separates AI dubbing from a generated voice
In conversations about AI several terms collapse into one, and those are exactly the terms that decide the scope of the work and the price. Let us take them apart:
- Speech synthesis (TTS, text to speech) — a model turning written text into speech; the result is a voice reading a line, with no relationship to the picture whatsoever.
- Voice cloning — building a model of one particular timbre from recordings of that person; it lets you speak in their voice, but says nothing about whether you may.
- Speech-to-speech — an actor performs the line at the microphone and the model swaps only the timbre; the performance, the timing and the emotion stay human.
- Lip-sync (the flaps) — the agreement between a line and the movement of the character’s mouth on screen; in dubbing it is what decides whether the viewer believes the character speaks their language.
- Prosody — the melody of a sentence: stress, pause, pace and the direction of the intonation; it is the difference between a line read out and a line spoken.
- Room tone — the quiet, particular hum of every interior; without it the voice has nothing to sit on and you can hear that it was pasted onto the picture.
- M&E (music and effects) — the full music and effects track with no dialogue on it; the basis of every language version.
- Stems — the separated layers of the finished mix (dialogue, music, effects), handed over with the mix for later use.
- Dialogue editing — laying every line out in time, cleaning it and fitting it into the rhythm and the acoustics of the scene.
A generator turns text into sound and its job ends there. Dubbing is script adaptation, casting a timbre, directing the performance, sync, dialogue editing and a mix — and, before any of that, an answer to the question of where the voice the character speaks in actually came from. The difference between the two is not the quality of the model; it is the number of stages somebody has to work through.
What the machine does, and what a person does
Dubbing and voice-over are trades with a craft of their own, and the people in them have good reasons to be wary of this technology. We share that wariness in principle, and our answer fits in one sentence: synthesis is a studio tool, not a replacement for an actor. We do not offer it to save the cost of a voice on a part that needs one, but on material that would otherwise never get a language version at all — a course refreshed every quarter, a dozen markets for one product video.
That said, refusing to see what these models genuinely do would be no more honest. So let us give them credit for what they handle:
- Timbre and vocal character — the model holds a consistent sound across the whole piece, with no fatigue and no differences between recording days.
- Baseline pace and correct reading — numbers, dates and long compound sentences come out evenly, which matters across hours of narration.
- Every language at once — the first pass of the track is produced in all versions in parallel, rather than one session after another.
- Scale — a long-form narration or a dozen language versions of one video stops being a question about booth availability.
- The cost of a change — one corrected line in a course updated every quarter no longer means calling an actor back in.
The rest sits with us. This is not a wish list; it is an inventory of the things you only hear once nobody has done them:
- Breaths and pauses — the model does not plan them, because it has neither lungs nor intent; dropped in mechanically they read as a tic, removed altogether they turn the character into a machine.
- Sentence stress — the emphasis lands where it statistically lands most often, not where the meaning of the line sits in this particular scene.
- Irony, whisper, effort, laughter — everything that depends on a gap between what is said and how it is said comes out flat.
- Stretching and trimming lines to the mouth — the model cannot see the face on screen and does not know when the character closes their lips.
- The acoustics of the room — a generation is made in nowhere: no reverb of the interior, no near and far perspective, no transitions between shots.
- The dynamics of crowd scenes — several voices generated separately settle into one flat plane instead of a foreground and a background.
- The collision with music and effects — only in the mix does it turn out that the dialogue disappears under the score, or steps out in front of the picture.
Sixteen steps between a model’s voice and finished dubbing
Below is the whole road from a file out of a model to a track you can send to air, with a note on who performs each stage. We publish this list because it is the most honest answer to the question of what you are paying for — and the simplest way for you to name what was missing from the cheap render that came back with notes on it.
- Spotting the material — human — we go through the material scene by scene: characters, on-screen and off-screen lines, the shots that need lip sync, lines that overlap. This is what determines the scope and the quote.
- Transcript check and speaker attribution — human — we correct proper names and specialist terminology, and mark the reactions, breaths and crowd sound that are not dialogue and should never reach the synthesis.
- Establishing the M&E track — human — we check whether the production has a music and effects version with no dialogue on it. If it does not, we plan the rebuilding of the backgrounds as a separate stage (step 15) and say so before quoting, not afterwards.
- Dialogue adaptation — human — this is not subtitle translation. The line has to sound like speech, fit inside the shot and inside a single breath of the character, not merely render the original sentence faithfully.
- Writing to the flaps (lip-sync) — human — in shots with a visible face we reorder the sentence and pick words that land on mouth openings and stressed syllables, so that the start and the end of the line agree with the picture.
- Client sign-off on the script — human — an agreed scope of notes on the text before anything is voiced. A change at this stage costs a minute; the same change after the mix costs a day.
- Casting the timbre and settling the source of the voice — human — we prepare samples and match the timbre to the character and to the target market. This is also where it is decided which of the two routes the voice comes from.
- Generating the language versions — machine — the model produces the first pass of the track, in every language at once. This is the one stage it performs on its own.
- Take selection and directing the performance — human at the machine — we run several passes at a line, steering the pace, the sentence stress and the emotion, then choose takes exactly as we would in a booth session with an actor.
- Pick-ups and ADR in the booth — human — the lines the model cannot carry — irony, whisper, a shout, laughter, effort — we record live to picture in our soundproof voice booth , with direction.
- Dialogue editing and sync to picture — human — we move, trim and stretch the lines until they fall into the rhythm of the scene and the movement of the mouth. It is the same work we do on dialogue recorded on set.
- Breaths, pauses and reactions — human — we put them in where the character genuinely takes air, and take them out where they sound mechanical. Without this stage a track is correct and dead at the same time.
- Placing the voice in the acoustics of the scene — human — room tone, near and far perspective, the reverb of the interior, transitions between shots. Without it the voice lies on top of the picture instead of standing inside the frame.
- Clean-up and repair — human at the machine — we remove synthesis artefacts, clicks at the joins and harsh sibilants, and even out uneven dynamics, with the same class of tools we use for audio restoration .
- Rebuilding the backgrounds and completing the M&E — human — where the music and effects track is missing or incomplete, we rebuild ambiences, Foley and effects under the new dialogue track, drawing on sound design and our sound library.
- Mix, loudness control and delivery — human — we seat the dialogue between music and effects, deliver the mix in stereo and 5.1, bring the material to the EBU R128 standard and export the stems and the M&E.
The model performs one of these sixteen steps. On two more we sit beside it and choose. The remaining thirteen are ordinary sound post-production — the same work we do on a film recorded entirely on set.
Whose voice is it
Before anything is voiced, it has to be clear where the voice came from. We treat that as a normal production stage, like spotting the material, rather than as terms and conditions bolted on at the end. For the synthesis itself we use commercially available, off-the-shelf tools — we have no model of our own and no voice library of our own — so there are two routes:
- A voice from the tool’s library — a timbre offered inside the tool we use. The right to use it comes from that tool’s licence, not from an agreement we hold: it is the tool’s provider who settles with the people whose voices it offers. Our job is to check before we start that the timbre comes from a lawful source and that the licence covers what the material is for — the channel, the languages and the markets.
- The voice of a particular person — an actor, a board member, a lecturer or an in-house presenter. We take that work only where we have their written consent in place, explicitly covering speech synthesis, with a defined scope and end date. It is a condition of taking the job, not a form to be completed along the way: without that document we do not start. It works well for internal communication and for training that is maintained over years.
On every project, which route the voice came from is documented — more on that below, under delivery.
What we will not do
There are jobs we will not take, and we would rather say so now than at the quoting stage:
- we will not clone the voice of a living actor or voice-over artist without their written consent covering speech synthesis;
- we will not recreate the voice of someone who has died without the consent of those entitled to give it;
- we will not build a voice that is confusingly similar to a recognisable person, including where it is formally not their voice;
- we will not replace a cast that holds a signed contract for roles on the production.
This is not a statement of belief, it is a condition of work. A soundtrack that leaves somebody with doubts about the rights to a voice comes back to the production faster than badly synced dialogue does.
When AI dubbing makes sense, and when we record a person
Not every piece of material suits this service, and we are not going to pretend otherwise. It makes sense for:
- e-learning and training localisation, especially courses updated every quarter, where every change would otherwise mean a fresh session;
- corporate and product material — presentations, instructions, internal communication;
- a dozen language versions of one video, where classic dubbing would mean a dozen separate castings;
- instructional and informational narration, where clarity matters more than a performance;
- background and secondary voices and crowd sound filling out the space of a scene.
We advise against it for leading roles in drama and animation, for brand advertising, for scenes that carry emotion, and always where there is no documented basis for the voice being used. In those cases we suggest recording the voice in the booth — with casting, actor direction and an ADR session. It is the more expensive service and we say so plainly, but on material meant to move an audience, saving money at this stage is what costs the most.
The productions we work on
This service fits:
- language versions of video — one piece of material, several or a dozen markets;
- e-learning and training — courses, onboarding, compliance material;
- corporate and product material — presentations, demos, user instructions;
- documentary and narration — off-screen commentary, archive material, institutional films;
- video channels and social media — content published regularly in several languages.
If your format is not on the list, write to us — once we have watched the material we will tell you whether this is the right service, or whether it is better to go towards classic voice recording .
What you get
Delivery is tangible and countable. When the work is finished we hand over:
- dialogue edited and synced to picture — ready for broadcast, not a raw generation;
- a mix in the agreed format — stereo and 5.1, depending on the distribution channel;
- stems — the separated layers of the mix (dialogue, music, effects) for further work;
- an M&E track — music and effects without dialogue, so the next language version does not mean rebuilding everything from scratch;
- versions compliant with loudness standards — EBU R128 and the technical requirements of the platform the material is going to;
- a documented origin for every voice — a record of which of the two routes each voice on the project came from;
- a statement marking the material as produced with the use of speech synthesis — prepared on request, or where the intended use of the material calls for it.
We treat the last two as part of the delivery rather than paperwork. They are what decides what you may do with the material in two years’ time, and in which markets — a question that comes up more and more often, because platforms and clients ask how a soundtrack was made.
Our studio
This service is not a layer of marketing over somebody else’s tool — it is our everyday post-production craft, with one of the stages performed by a model.
Olga Pasternak works in dialogue editing, ADR and sound editing. She began her career as an assistant to the Oscar-winning sound designer Nicolas Becker, on the film Kursk. As supervising dialogue editor she worked on the Oscar-winning “The Substance”; she has edited dialogue and ADR on the series Lupin (Netflix), Around the World in 80 Days and Infiniti, and also worked on Une Année Difficile. Yohan Boisgontier handles sound design, Foley and the mix — on the film Triumph of the Heart he served, together with Olga, as supervising sound editor.
We work in a 5.1 editing and mix room with controlled acoustics, on Pro Tools Ultimate. Pick-ups and ADR are recorded in a soundproof voice booth with a Neumann TLM 102 microphone, with an HD screen, a camera for direction and a talkback linked to the mix control room. For cleaning tracks we use iZotope RX-class restoration tools, and we build backgrounds and effects from a 4 TB sound library. We record and direct in Polish, English, French and Spanish.
Our dubbing studio is in Warsaw, in the Żoliborz district, at ul. Śmiała 42. Playbacks and direction happen on site or remotely — you listen to the session and give notes without leaving your own office.
Why we quote every project individually instead of publishing a price list
We do not have a dubbing price list, and we treat that as an advantage rather than a gap. A price list would have to assume the most labour-intensive scenario — and you would then be paying for stages your material does not involve at all. We quote exactly those of the sixteen steps your project needs, and we say plainly which ones fall away.
What drives the workload on dubbing services is mainly:
- the length of the material and the density of dialogue — an hour of a lecture with a single speaker is different work from twenty minutes of scenes with several characters talking at once;
- the number of language versions — in video localization for several markets, a large part of the work (track preparation, mix template, matching the backgrounds) is done once and reused in every version, so the next language costs less than the first;
- the level of synchronisation required — off-screen narration, matching the start and end of each line, and full lip sync are three different amounts of dialogue editing;
- whether an M&E (international) track exists — with music and effects delivered separately we can drop the new dialogue straight in; without it, the track has to be rebuilt first: dialogue cut out of the mix, backgrounds filled in, missing effects and Foley recorded;
- the state of the source material — a script and a time-coded dialogue list shorten the work; without them we start by transcribing the dialogue from the soundtrack;
- how much has to be recorded in the booth — individual lines, post-sync and ADR with an actor are a separate stage, and in some projects it does not occur at all;
- whether the material will be updated — for training and product videos we build the project so that the next update means adding a section, not re-recording everything.
To get a quote we only need the video (or a review link), the list of target languages and your deadline. If you have a script, a dialogue list or an M&E track, send those too — then we give you a specific figure and schedule rather than a range.
Order AI dubbing and language versions
We do not work from a fixed price list — send us the material and we will prepare a quote and a schedule, setting out which of the sixteen steps your project needs and which fall away. Write to mail@allinsound.studio or use the contact page , where you will also find direct phone numbers for Olga and Yohan. Tell us how many language versions you are planning, whether an M&E track exists and whether the material will be updated later. We also work with clients outside Poland; reviews and playbacks can be run remotely.
We deliver AI dubbing as a standalone service or as part of full post-production — together with voice-over and voice recording , sound design , mix and mastering and audio restoration . See also the full studio offer .
Want to read up first? We explain what dubbing is , what ADR and post-sync are and what sound post-production is .
Equipment
- Pro Tools Ultimate — dialogue editing, sync and delivery
- 5.1 editing and mix room with controlled acoustics
- Soundproof voice booth with a Neumann TLM 102 — pick-ups and ADR
- HD screen and camera for direction, talkback to the mix control room
- iZotope RX-class restoration tools and a 4 TB sound library
- Language versions in PL, EN, FR, ES — stereo and 5.1, stems and M&E, EBU R128
Frequently asked questions
How is AI dubbing different from a voice out of a generator?
A generator turns text into sound and its job ends there. Dubbing means dialogue adaptation, casting the timbre, directing the performance, sync to picture, dialogue editing, placing the voice in the acoustics of the scene and mixing it against music and effects. The model handles one of the sixteen steps we list on this page. On two more we sit beside it and choose; the remaining thirteen are ordinary sound post-production, done exactly as on a film recorded on set.
Does this take work away from actors?
Not the way we use it. We do not offer synthesis to save the cost of a voice on a part that needs one: on drama, animation and advertising we record a person in the booth, with casting and direction. Synthesis goes where the material would otherwise never get a language version at all — a course refreshed every quarter, a dozen markets for one product video. And where the voice of a particular person is involved, their written consent covering speech synthesis is a condition of us taking the job.
Is AI dubbing suitable for a feature film?
For leading roles and scenes that carry emotion we advise against it. The model works well in informational narration, e-learning, corporate material, and background and secondary voices. If the material has to carry a performance, we suggest recording the voice in the booth with actor direction. That is a different service and a different budget, and we say so before quoting.
Where does the voice come from, and is it legal?
For the synthesis itself we use commercially available, off-the-shelf tools — we have no model of our own and no voice library of our own. That leaves two routes. The first is a voice from the library of the tool we use: the right to use it comes from that tool’s licence, and it is the tool’s provider who settles with the people whose voices it offers; we check before we start that the licence covers what the material is for. The second is the voice of a particular person — an actor, a board member, a lecturer, an in-house presenter — and in that case their written consent, explicitly covering speech synthesis, with a defined scope and end date, is a condition of us taking the job. On every project, which route the voice came from is documented.
Can you clone the voice of a particular actor or voice-over artist?
Only where that person has given us their written consent covering the synthesis of their voice, and only within the scope that consent allows. We do not clone a voice without the owner’s consent, we do not recreate the voice of someone who has died without the consent of those entitled to give it, and we do not build a voice that is confusingly similar to a recognisable person. This is a condition we do not waive.
Will I get an M&E track and stems?
Yes. Alongside the finished mix we deliver an M&E track — music and effects without dialogue — and stems, meaning the separated layers of the mix. That way the next language version can be produced later without rebuilding the whole soundtrack. If the production has no M&E track, we plan the rebuilding of the backgrounds as a separate stage of the work.
Which languages do you produce versions in?
We record and direct in Polish, English, French and Spanish — in those languages we can judge whether the phrasing sounds natural and run pick-ups in the booth ourselves. For wider language sets we agree the scope of verification individually: which versions go through a full listening check on our side, and which need a native speaker on yours.
Will the dialogue match the lip movement?
In shots where the face is visible we write the lines to the flaps: we reorder the sentence, choose words that fall on mouth openings and stressed syllables, and in the edit we move, trim and stretch the lines. In off-screen narration and training material, sync is about events on screen rather than lips, so there is less of this work. We assess the scope once we have watched the material.
What do I need to send, and can we work remotely?
Ideally a video with a burned-in timecode window, a transcript or the original script, the M&E track if one exists, and a character list with descriptions. It helps to know whether the material will be updated and which markets it is going to. We also work with clients outside Poland; reviews, direction of pick-ups and playbacks can be run remotely.
How much do AI dubbing services cost?
We do not work from a fixed price list. The cost depends on the length of the material, the number of language versions, whether lip sync is involved, whether an M&E track exists and how many lines have to be recorded live. Send us the material and we will prepare a quote and a schedule that shows which stages you are paying for.
Will the material be marked as produced with AI?
Yes, if you want it or if the intended use of the material calls for it — we then prepare a statement that speech synthesis was used in producing the soundtrack. Either way, on every project we document which route each voice came from. That is what tells you what you may do with the material later, and in which markets.



