Transcriptor
The Transcriber is an all-in-one solution designed to transcribe your calls and other voice interactions into text. The workflow comprises several programmable steps:
- Data Gathering & Parameter Ingestion
- Correlation & Independent Track Analysis
- Crosstalk Detection & Suppression
- Audio Normalization
- Distortion Detection
- Diarization
- Overlay Detection & Mitigation
- Transcription
- Transcription Confidence Redo
- Automated Tagging
Data Gathering & Parameter Ingestion
In this step, Contacter gathers the relevant information required to identify the call, such as:
- Line In: The called number or target service.
- Agent: The Customer Service Representative (CSR) who handled the call.
- Customer: The calling customer.
- Date: The exact date and time of the call.
These fields are mandatory. Contacter will attempt to automatically extract these values from the provided filename. You can complement this data with explicit parameters or define the filename format to guide the parser.
The Transcriber accepts the following execution parameters:
filein– Specific file to process or a folder directory to process (iterating through all files sequentially).pathout– The target directory where the output CTT file will be saved.-lor--line– Forces a specific Line In value (optional).-aor--agent– Forces a specific Agent value (optional).-ror--customer– Forces a specific Customer value (optional).-dor--date– Forces a specific Date value; limited to single-file execution (optional).-tor--timezone– Provides or overrides the execution timezone (optional).-sor--separator– Indicates the field separator used in the filename; attempts to auto-detect if omitted (optional).-for--format– Indicates the field order within the filename, e.g.,lacd(optional).-eor--move– Automatically moves the original audio file to a designated folder upon successful processing (optional).-vor--verbose– Sets the verbosity level, ranging from 0 (minimal) to 4 (maximum) (optional).-por--project– Defines a project name to segment and group calls logically (optional).-gor--tag– Defines an alternative to the use of tag.yaml Fiele*(optional)*.
Correlation & Independent Track Analysis
This step detects correlation between audio tracks to identify stereo configurations or strongly correlated tracks, producing a list of genuinely independent channels. This prevents the system from generating duplicate transcriptions for highly correlated tracks.
Contacter distinguishes between track correlation and track crosstalk. Crosstalk refers to signal bleeding between otherwise independent channels. To ensure accuracy, correlation detection performs crosstalk suppression prior to calculating the final correlation values.
# Multi-track processing configuration
# docheck_correlation: Set to True to verify if tracks are correlated (e.g., stereo layouts)
# correlation_window: Window size (in seconds) used for correlation analysis
# silence_limit: Noise floor/amplitude threshold for crosstalk analysis
# pearsoncorrelation_limit: Statistical threshold defining a valid correlation
# pearsoncorrelation_mix: Required percentage of correlated windows to classify tracks as correlated
# crosstalk_limit: Power difference threshold used to identify crosstalk between tracks
MULTIPLE_CHANNEL:
docheck_correlation: True
correlation_window: 5
silence_limit: 1e-4
pearsoncorrelation_limit: 0.65
pearsoncorrelation_mix: 0.65
crosstalk_limit: 0.15
The Transcriber detects and mixes correlated tracks (such as stereo split layouts) to avoid redundant transcriptions via the following pipeline:
- Compares every audio track against every other available track.
- Splits each track into blocks defined by
correlation_window(in seconds). - Calculates the correlation coefficient for each block using the Pearson Formula:
- If any track amplitude falls below
silence_limit, it is treated as uncorrelated noise. - If one track's power drops below the
crosstalk_limitrelative to the other, it is treated as uncorrelated. - If neither condition is met and the Pearson correlation exceeds
pearsoncorrelation_limit, the block is flagged as correlated.
- If any track amplitude falls below
- If the overall percentage of correlated windows exceeds
pearsoncorrelation_mix, the tracks are classified as Strongly Correlated and are automatically mixed.
Crosstalk Detection & Suppression
This step isolates crosstalk between independent channels and suppresses it to eliminate duplicate transcriptions.
# Crosstalk detection and mitigation variables
# silence_limit: Amplitude threshold defining silence or background noise floor
# pearsoncorrelation_limit: Value threshold defining a statistically correlated signal
# docrosstalk_correction: Set to True to enable automated crosstalk correction and gating
# crosstalk_limit: Power difference threshold between channels indicating a crosstalk leak
# crosstalk_window: Block size (in milliseconds) used to evaluate crosstalk occurrences
MULTIPLE_CHANNEL:
silence_limit: 1e-4
pearsoncorrelation_limit: 0.65
docrosstalk_correction: True
crosstalk_limit: 0.15
crosstalk_window: 100
The Transcriber executes crosstalk detection and suppression across independent tracks (typically after the Correlation Mix phase) using the following logic:
- Every track is segmented into short blocks defined by
crosstalk_window(in milliseconds). - Windows are overlapped with preceding and succeeding blocks to smooth out suppression artifacts.
- If the tracks exhibit correlation (exceeding
pearsoncorrelation_limit) and neither track is fully silent (abovesilence_limit), but one channel's power is significantly lower than the other (withincrosstalk_limit), the weaker signal is classified as a leak and suppressed. This suppression process utilizes a mathematical Hann Window function to maintain audio smoothness.
Audio Normalization
This step normalizes the audio amplitude of each independent track to optimize ingestion for subsequent diarization and transcription engines.
# Audio Amplitude Level Normalization Layout
# dopre: Enables or disables global pre-level normalization (default: True)
# pre_gain: Applies a fixed pre-level gain coefficient (set to 0 to disable)
# pre_quantile: Sets a specific quantile for peak estimation (set to 0 to disable)
# pre_peak: Target peak level for normalization (relative to absolute 1.0 full scale)
# pre_points: List of dynamic quantile thresholds tested by the optimization algorithm
# pre_change: Evaluation delta multiplier between dynamic quantile points
# dotranscription: Enables or disables automated phrase-by-phrase normalization
# transcription_gain: Applies a fixed gain to individual phrases during transcription
# transcription_quantile: Quantile target used to capture useful peaks per phrase
# transcription_peak: Target target peak level for transcription phrase normalization
LEVEL_NORMALIZE:
dopre: True
pre_gain: 0
pre_quantile: 0
pre_peak: 0.7
pre_points: [0.95, 0.97, 0.98, 0.99, 0.995, 0.999]
pre_change: 1.4
dotranscription: True
transcription_gain: 0
transcription_quantile: 0.99
transcription_peak: 0.9
You can apply a global gain adjustment across an entire independent track by enabling dopre. To prevent digital clipping and audio distortion from excessive gain, the normalization logic targets the loudest segments of the audio.
However, because sporadic, non-voice audio spikes can skew traditional peak normalization, the system evaluates a specific statistical quantile (such as 97%, 98%, or 99%) to determine the useful peak rather than using the absolute mathematical peak.
You can configure this gain using three mutually exclusive strategies:
- Fixed Gain: Set a direct gain value via
pre_gain. This is typically a value above 1, but is not recommended unless the original recording volumes are highly standardized. - Static Quantile: If
pre_gainis set to 0 or omitted, the second strategy uses a fixed quantile target (pre_quantile) to identify the useful peak. The engine calculates the required gain to bring that useful peak up to thepre_peaktarget level. - Dynamic Quantile Optimization: If both
pre_gainandpre_quantileare 0 or omitted, the Transcriber automatically runs an optimization loop across the array defined inpre_points. The algorithm evaluates the useful peaks for each sequential quantile step and selects the final quantile where the ratio of the previous peak to the current peak exceeds thepre_changethreshold. This fine-tuning process effectively isolates and ignores non-voice spikes.
Once speaker diarization is complete, the transcription engine processes the file segmentally, phrase by phrase. You can choose to individually normalize each phrase to maximize recognition accuracy. Enabling dotranscription applies a fixed modifier via transcription_gain. If that is set to 0, it dynamically calculates a phrase-level adjustment using the transcription_quantile to hit the target transcription_peak.
Distortion Detection
This step actively monitors, flags, and logs digital audio distortion within the file metadata.
# Digital Saturation and Distortion Tracking
# distortion_peak: Amplitude threshold above which a sample is flagged as distorted
# distortion_samples: Minimum number of consecutive clipped samples required to log distortion
LEVEL_NORMALIZE:
distortion_peak: 0.98
distortion_samples: 10
Diarization
This step executes speaker diarization, identifying each unique speaker and mapping them to their respective phrases. The diarization engine is powered by Pyannote modules. We have two versions to perform diarization.
# Diarization using transformed ONNX files. These are external produced and full credit is due for the authors
# directory: directory for the ONNX models and author credits
# segmentation: segmentation ONNX file name
# embedding: embedding ONNX file name
# window_step: start 10 seconds analisys every N seconds
# min_voice_probability: if within a speaker phrase another speaker is 0.N strengh we create a phrase for the second speaker
# min_duration: ignore blocks below N seconds
# silence_phrase_break: break a setence above N seconds
# min_alone_voice_embedding: if a speker has less than 0.N time speaking alone we dont separate the speaker in that windows
# min_voice_windows: 0.N samples must agree that in a specific windows the speaker was active
# min_final_duration: below N seconds ignpre phrases
# providers: hardware ro choose from ["CUDAExecutionProvider", "CPUExecutionProvider"]
# clustering : speaker clustering ensibility
# n_clusters: null
# metric: "cosine"
# linkage: "average"
# distance_threshold: 0.58 (select the sensibility to avoid ghost speakers or mixexing different speaker)
DIARIZATION_ONNX:
directory: "/var/opt/contacter/ai_models/onnx"
segmentation: "segmentation.onnx"
embedding: "embedding.onnx"
window_step: 0.5
min_voice_probability: 0.5
min_duration: 0.170
silence_phrase_break: 1.5
min_alone_voice_embedding: 0.2
min_voice_windows: 0.5
min_final_duration: 0.3
providers: ["CUDAExecutionProvider", "CPUExecutionProvider"]
clustering :
n_clusters: null
metric: "cosine"
linkage: "average"
distance_threshold: 0.58
Diarization is performed utilizing the Pyannote ONNX versions hosted on Hugging Face (https://huggingface.co/pyannote). The instalation script will copy the ONNX files and respective credits to
a selected directory.
# Diarization using Pyannote model and authentication token layout (in alpha testing. available very soon)
# \$key: Reference to sensitive variables defined in keys.yaml
DIARIZATION:
hp_token: "$key:pyannote_hp_token"
model: "pyannote/speaker-diarization-3.1"
segmentation:
# min_duration_off: 0.1 # Minimum Silence to close a Phrase
clustering:
# method: centroid
# threshold: 0.82 # Increase if ghost Speaker
execution:
# min_speakers: 1
# max_speakers: 10
Diarization is performed utilizing the Pyannote ecosystem hosted on Hugging Face (https://huggingface.co/pyannote). To use Pyannote's diarization pipelines, you must register a free account on Hugging Face and manually accept the user terms for both the pyannote/speaker-diarization and pyannote/segmentation models. Once accepted, generate a User Access Token to authenticate your environment. This token should be securely mapped inside your keys.yaml file.
Overlap Detection & Mitigation
This step detects overlapping audio segments from different speakers and mitigates them to avoid misattributing spoken phrases. Contacter isolates and splits overlapping phrases based on two scenarios:
- Partial (Single) Overlap
- Total Overlap
# Speaker Overlap Correction and Gating
# dooverlap: Enables or disables automated overlap filtering (default: True)
# allowed: Maximum allowed overlap duration (in seconds) before triggers execute
# apply: Small crossfade or padding adjustment (in seconds) applied to cut boundaries
# segmentmin: Minimum duration threshold (in seconds) required to retain a segment
OVERLAP_CORRECTION:
dooverlap: True
allowed: 0.2
apply: 0.1
segmentmin: 0.5
The Transcriber flags an overlap event if two distinct phrases intersect for longer than the allowed threshold.
- In a Partial Overlap scenario, the first phrase is trimmed exactly at the start boundary of the succeeding phrase.
- In a Total Overlap scenario, the underlying phrase is split into two distinct segments, sandwiching the interrupting speaker's phrase in the middle.
A small buffer defined by apply (in seconds) is added to the boundaries to smooth out transitions. Any resulting audio fragments shorter than the segmentmin limit are automatically discarded as negligible noise.
Transcription
This step handles the automated speech-to-text conversion for each audio segment. Transcription tasks are executed using the faster-whisper backend.
# WhisperModel core engine configurations
# Custom parameters appended here will add to or overwrite system defaults
# Supports all standard Faster-Whisper parameters plus RESET and model_size
# RESET: Set to True to completely discard all built-in system defaults
# model_size: Map this variable to the first unnamed positional parameter of the model
# cpu_threads: Values < 1 map to a percentage of total host cores; values >= 1 clamp to absolute cores
WHISPER_MODEL:
#RESET: True
model_size: "large-v3-turbo"
cpu_threads: 0.5
The Transcriber prioritizes NVIDIA CUDA hardware acceleration whenever a compatible GPU is available. If no GPU environment is detected, it falls back to multi-threaded CPU execution.
When CUDA acceleration is available, the system enforces the following environment defaults:
device: "cuda"model_size: "large-v3-turbo" (if system VRAM > 6GB) or "medium" (if VRAM <= 6GB)compute_type: "float16" (if system VRAM > 4GB) or "int8_float16" (if VRAM <= 4GB)
When CPU fallback occurs, the system operates under these defaults:
device: "cpu"model_size: "small"compute_type: "int8"cpu_threads:int(cpucount/2)(if the host machine has more than 1 core) or 1
You can alter these settings or completely wipe the environment defaults by declaring RESET: True. The cpu_threads variable is highly flexible: declaring an integer (cpu_threads >= 1) locks execution to that exact thread count, while declaring a float (0 < cpu_threads < 1) allocates a precise percentage of the host machine's total logical processing threads (e.g., 0.5 allocates 50% of available CPU cores).
# WhisperTranscribe execution variables
# Custom parameters declared here will append to or overwrite system defaults
# Supports all standard Faster-Whisper transcribe options (except initial_prompt)
# RESET: Set to True to discard built-in default values
# initial_prompt: Handled dynamically via the WHISPER_PROMPT section below
WHISPER_TRANSCRIBE:
#RESET: True
beam_size: 5
best_of: 5
repetition_penalty: 1.2
word_timestamps: True
vad_filter: True
vad_parameters: {"min_speech_duration_ms": 300}
The Transcriber relies on the high-performance faster-whisper library for its transcription core. The variables listed above mirror the standard developer options found within that library and represent the platform's out-of-the-box defaults. You can tweak any entry or wipe out the defaults entirely by passing RESET: True.
# WhisperTranscribe initial_prompt dynamic payload builder
# glossary: Custom vocabulary, brand acronyms, or specific technical terms
# style: Contextual hint defining conversational tone or speech patterns
# language: Explicit target language configuration to prevent model drifting
# continue: Number of tail words captured from a split segment to prime the next prompt
# contexclude: Set to True to omit static glossary/style tokens when continuing a split segment
WHISPER_PROMPT:
glossary: ""
style: ""
language: ""
continue: 15
contexclude: True
This parameters block defines how the platform dynamically shapes the initial_prompt payload passed to Faster-Whisper. The fields glossary, style, and language are automatically concatenated into the prompt if populated.
The variables continue and contexclude manage contextual steering when a continuous speech segment by a single speaker is chopped up due to an overlap event. To maintain semantic consistency across the split, the subsequent transcription segment can be primed with up to continue words from the end of the previous segment (set to 0 to disable).
Enabling contexclude tells the prompt manager to temporarily omit the static glossary, style, and language rules when a text continuation is active, preventing prompt pollution and model confusion from overloaded context windows.
Transcription Redo
This step identifies phrases with low language-detection confidence (applicable only when a static target language is not forced during the initial transcription phase). You can define specific logical conditions that trigger an automated re-transcription of those phrases under a strictly forced-language condition.
WHISPER_REDO:
# The automated re-evaluation loop is executed only if set to True
redo: True
# These Cequation lines are evaluated for every individual dialogue entry (phrase) in the call.
# The objective is to return True if a transcription REDO should be initiated.
# 'dialog' acts as an array of all dialogue phrases captured during the interaction.
# Use '[my]' to reference the index/key of THIS specific current phrase.
# 'langprob' is the metadata field defining the statistical language probability score.
# In Polish Notation, '0.91 <' evaluates to True if the language probability falls below 91%.
# The assignment key below does not start with a '$'. This indicates it is not a system CTT variable,
# but a local temporary variable named 'dialog-langprob-is-small'.
"dialog-langprob-is-small": "$dialog.[my].langprob 0.91 <"
# This Cequation compares the current phrase language with the speaker's primary language.
# 'speaker' is a dictionary of all call participants; '[my]' targets the speaker of the current phrase.
# 'language' is a sub-dictionary tracking all languages spoken by this speaker during the call.
# Each language key maps to its specific spoken duration.
# The range modifier 'duration:-1' extracts the sub-dictionary element with the longest duration (-1 indicates the last sorted element).
# The suffix '.$key' extracts the string key name of that sub-dictionary rather than its inner content.
# The current phrase language is compared against this primary language; if they differ, True is assigned to the local variable.
"dialog-language!=speaker-max-language": "$dialog.[my].language $speaker.[my].language.duration:-1.$key !="
# Here we retrieve the duration of the speaker's primary language and divide it by their total spoken duration.
# If the speaker uses their primary language more than 80% of the time, the condition evaluates to True.
"speaker-max-language>80": "$speaker.[my].language.duration:-1.duration $speaker.[my].duration / 100 * 80 >"
# Finally, the core 'condition' variable dictates whether the dialogue phrase undergoes a REDO.
# Prefixing a variable name with a dot (e.g., '.dialog-langprob-is-small') invokes the local temporary variable.
# The system evaluates the three preceding local variables and performs a sequential 'and' operation across them.
# Logical boolean operations are executed, and the final 'and' joins the resulting boolean with the speaker's primary language string.
# Performing an 'and' between a Boolean and a String returns the string value if True, or an empty string ("") if False.
# The 'condition' key returns either a target language code or a Boolean. If it returns a language string,
# that specific language is forced during the re-transcription process. If it returns True, a REDO is performed
# using the speaker's most spoken language. If it returns False or "", the phrase is skipped.
"condition": ".dialog-langprob-is-small .dialog-language!=speaker-max-language .speaker-max-language>80 and and $speaker.[my].language.duration:-1.$key and"
# In this pipeline example, a transcription REDO with a forced language condition executes if the language probability
# drops below 0.91 AND the current segment language differs from the speaker's primary spoken language.
# If the resolved target language matches the existing phrase language, the REDO operation is automatically bypassed.
# You can customize these threshold conditions freely to force targeted re-transcriptions.
The REDO configuration consists of a sequential series of Cequations used to determine if a transcript segment requires reprocessing. The final condition variable acts as the execution gate. The condition key can return an explicit language string, a simple True value (which forces the re-transcription to fall back onto the speaker's primary language), or False/"" to indicate that no REDO operation should be performed. For more advanced configurations, please consult the Cequation core documentation.
Tagging
This step handles the automated tagging of the call, individual speakers, and distinct dialogue phrases. These tags help downstream AI models better understand contextual nuances and behavioral metadata that are not explicitly captured within the raw transcript text. Tagging configurations are driven entirely by Cequation logic. For a deep dive into expression syntax, refer to the core Cequation documentation.
# CALL-LEVEL TAGGING INDEX
TAG_CALL:
dotag: True
$mainspeaker: "$speaker.duration:-1.$key"
$mainlanguage: "$call.language.duration:-1.$key"
# SPEAKER-LEVEL TAGGING INDEX
TAG_SPEAKER:
dotag: True
$mainlanguage: "$speaker.[my].language.duration:-1.$key"
$mainlanguagevalue: "$speaker.[my].language.duration:-1.duration $speaker.[my].duration / 100 * round.0"
$RECORDED_VOICE: "$speaker.[my].logprob -0.28 > $speaker.[my].language.langprob:wavg 0.98 > $speaker.[my].nospeechprob 0.05 < and and"
$MUSIC: "$speaker.[my].nospeechprob 0.6 > $speaker.[my].volumecorr 0.02 > $speaker.[my].comratio 2.4 > and and"
"ivr_val": "$dialogspk.text:join text.in.marq.mark.apoio.avari.fatur.aces.intern.velocid"
$IVR: ".ivr_val $speaker.[my].turns 2 / >"
"announce_val": "$dialogspk.text:join text.in.gpdr.comercia.grava.lega.prov.contrat.contact.informa"
$ANNOUNCE: ".announce_val $speaker.[my].turns 2 / >"
# PHRASE/DIALOGUE-LEVEL TAGGING INDEX
TAG_DIALOG:
dotag: True
$UNCERTAIN_AUDIO: "$dialog.[my].logprob -0.7 < $dialog.[my].nospeechprob 0.4 < and"
$BACKGROUND_NOISE: "$dialog.[my].nospeechprob 0.6 > $dialog.[my].volumecorr 0.02 > $dialog.[my].comratio 2.4 < $dialog.[my].logprob -1.0 < and and and"
$WEAK_SIGNAL: "$dialog.[my].volumecorr 0.01 < $dialog.[my].nospeechprob 0.5 < and"
$AUDIO_CUTS: "$dialog.[my].duration 0.8 < $dialog.[my].logprob -1.2 < and"
$DISTORTION: "$dialog.[my].distortion $dialog.[my].distortion 1 >"
$UNRECOGNIZED: "$dialog.[my].logprob -1.8 < $dialog.[my].nospeechprob 0.9 > and"
dtmf_val1: "$dialog.[my].nospeechprob 0.8 > $dialog.[my].duration 0.1 > $dialog.[my].duration 0.5 < and and"
dtmf_val2: "$dialog.[my].text text.digit $dialog.[my].duration 0.7 < and"
$DTMF: ".dtmf_val1 .dtmf_val2 or"
long_silence_val: "$dialog.[my].start $dialog.[my]-1:0:num.end - round.1"
$LONG_SILENCE: ".long_silence_val postfix.s .long_silence_val 4 >"
high_tone_val: "$dialog.[my].volumecorr $speaker.[my].volumecorr / 100 * round.0 100 -"
$HIGH_TONE: ".high_tone_val prefix.+ postfix.% .high_tone_val 40 > $dialog.[my].nospeechprob 0.2 < and"
fast_speak_val: "$dialog.[my].wordnumber $dialog.[my].duration / round.1"
$FAST_SPEAK: ".fast_speak_val postfix.wps .fast_speak_val 4.5 >"
whisper_val: "100 $dialog.[my].volumecorr $speaker.[my].volumecorr / 100 * round.0 -"
$WHISPER: ".whisper_val prefix.- postfix.% .whisper_val 30 > $dialog.[my].nospeechprob 0.2 < and"
overlap_val: "$dialog.[my].overlap.duration:sum:0:num $dialog.[my].duration / 100 * round.1"
$OVERLAP: ".overlap_val postfix.% .overlap_val 25 >"
The TAG blocks contain a structured series of Cequations that define the evaluation boundaries for applying metadata attributes. Within these scopes, only lines that begin with a $ prefix represent actual system tags; lines omitted of this prefix serve as local auxiliary variables.The configuration isolates parameters into three distinct structural scopes: call-level indicators (TAG_CALL), speaker-level signatures (TAG_SPEAKER), and individual phrase segments (TAG_DIALOG).To illustrate this runtime behavior, consider the execution path of the long silence detection rule:
- long_silence_val: "$dialog.[my].start $dialog.[my]-1:0:num.end - round.1"
- $LONG_SILENCE: ".long_silence_val postfix.s .long_silence_val 4 >"
The first expression instantiates an auxiliary variable that captures the timestamp tracking the start of the current phrase segment ($dialog.[my].start) and the timestamp tracking the conclusion of the immediately preceding dialogue entry ($dialog.[my]-1:0:num.end). The engine subtracts the values and rounds the resulting delta to a single decimal place (- round.1).Because a preceding phrase does not exist when parsing the initial dialogue index, passing the :0:num fallback modifier instructs the runtime parser to return a numerical 0 value rather than throwing a null exception. This final delta is stored inside the local tracker long_silence_val.The succeeding Cequation appends a "s" unit indicator to the value string and evaluates whether the contents of long_silence_val exceed 4 seconds. If the statement yields a logical true, the system appends the structural metadata response to the schema (e.g., LONG_SILENCE=4.3s). Conversely, if the boolean evaluation yields a false state, the tag allocation is skipped.