Skip to main content
Smart Diarization detects who speaks when in a mixed recording, lets you review the detected speakers, then creates a new sequence with one audio track per speaker. It is useful when you received a single mixed audio file but need separate speaker tracks for editing, mixing, captions, or multicam preparation.

Quick Start

1

Open the source sequence

Make the sequence containing the mixed conversation active.
2

Choose an export folder (optional)

Select a permanent folder next to your project, or leave the field empty to use the Premiere Copilot folder in Documents.
3

Detect speakers

Click Detect Speakers and wait for the audio export and speaker detection to finish.
4

Review the speakers

Check the number of detected segments and speaking time. Click a speaker to jump to their first line.
5

Split the speakers

Click Split Speakers to create the new speaker sequence.
The exported audio file becomes the media used by the new sequence. Do not delete or move it after the split unless you relink it in your editing software.

What the result contains

The review displays each detected speaker with:
  • A speaker name.
  • The destination audio track.
  • The number of speaking segments.
  • The total speaking time.
  • A shortcut to the first detected line.
When applied, the tool imports the exported audio and builds a new sequence in which each speaker is heard only on their assigned track. Your original sequence remains available.
Speaker detection separates voices by identity, not by microphone name. If two people sound very similar or speak over each other frequently, review the result carefully.

Choosing the export folder

Use a folder that will stay with the project. The default Premiere Copilot folder is suitable for local work, while a dedicated media folder next to the project is safer for archiving or handing the project to another editor. Saved presets remember your preferred export folder.

Best source material

Smart Diarization works best when:
  • Each person has a distinct voice.
  • Dialogue is louder than music and background noise.
  • Speakers do not constantly talk over one another.
  • The sequence contains continuous, online audio.

Common pitfalls

  • Music under the conversation can be interpreted as part of the mix and reduce separation quality.
  • Heavy overlap between speakers makes the boundary between voices uncertain.
  • Deleting the exported audio takes the new sequence offline.
  • Applying before review can create an unnecessary sequence when the detected speaker count is wrong.
  • Editing during processing can invalidate the timing used to rebuild the tracks.
For the cleanest result, duplicate the source sequence, mute music and sound effects, run Smart Diarization on dialogue only, then use the separated sequence in your final edit.