Skip to main content

Ainter

The Ainter utility utilizes Artificial Intelligence to enrich and augment the information contained within CTT files. Following the initial transcription phase, you can execute several programmable AI operations, including:

  • Identifying distinct speaker profiles.
  • Generating comprehensive call summaries.
  • Extracting structured lists of discussed issues and resolutions.
  • Executing any other custom AI-driven analysis required by your workflow.

Ainter provides the following core features:

  • Dynamic Prompt Engineering: Generates complex LLM prompts by combining version-controlled components with real-time data extracted from the CTT file.
  • LLM Orchestration: Invokes selected Large Language Models (LLMs) and securely logs execution metadata and results.
  • Structured Parsing: Analyzes JSON payloads returned by the LLM and seamlessly maps the structured data back into the target CTT file.

Parameters​

The Ainter utility accepts the following execution parameters:

  • filein – The target CTT file to process.
  • pathout – The directory where the newly enriched CTT file will be generated.
  • action – The specific identifier of the action block to execute.
  • -e or --move – Automatically moves the original CTT file to a designated folder upon successful processing (optional).
  • -v or --verbose – Sets the verbosity level, ranging from 0 (minimal) to 4 (maximum) (optional).
  • -t or --testprompt – Generates and displays the prompt payload for testing purposes without invoking the remote LLM (optional).

Ainter YAML Configuration Files​

All specifications regarding prompt building, model hyper-parameters, and JSON response parsing are defined within ainter_*.yaml configuration files. These contextual operations are structured as Actions. At runtime, the tool automatically discovers and parses all ainter_*.yaml files located within the current working directory. You can organize your logic across multiple files as preferred; certain prompt components can be dynamically shared across different actions.

An ainter_*.yaml configuration layout isolates logic into five core component classes:

  • model: Specifies the targeted LLM models, API connection strings, and endpoint configurations. You can define multiple provider blocks and bind them selectively inside your actions.
  • system: Defines the granular block components used to assemble the SYSTEM prompt. It supports raw text mixed with dynamic variable tokens pulled from the underlying CTT data model.
  • user: Defines the granular block components used to assemble the USER prompt. It supports raw text mixed with dynamic variable tokens pulled from the underlying CTT data model.
  • output: Defines the validation schemas and parsing paths used to process incoming structured JSON outputs from the LLM.
  • action: Functions as the orchestration block for execution. Within an action, you map which model block to load, which system and user prompt arrays to concatenate, and which output schema to evaluate. It also exposes custom loop iterations designed to process data iteratively across separate speakers or individual dialogue entries.

Component Versioning & Tags​

To maintain strict tracking across environments, every component block features a dedicated Name and Version. You can seamlessly introduce new blocks or publish iterative versions under an existing name block. A version identifier can be declared using an absolute numeric index (e.g., 1.1) or a tracking environment alias string (e.g., stable).

When binding components inside an action block, you declare the target component using a name:version formatting layout:

  • Referencing identify:1.1 binds the explicit version 1.1 of the identify block.
  • Referencing identify:stable instructs the runtime engine to dynamically fetch the highest indexed version tagged as stable within that namespace block.

Components and versions can be structured across centralized files or scattered across action-specific documents depending on your architectural preferences.


Model Component​

The Model component manages authentication variables, endpoints, and hyper-parameters required to interface with local or cloud-hosted LLM endpoints. The configuration example below details two distinct endpoint configurations: an on-premise infrastructure instance named llama3.2-internal and a cloud gateway instance named openrouterautofree. Both blocks track version 0.1 and are marked with a stable alias tag. Contacter currently supports native gateway connectors for Ollama and OpenRouter (OpenAI-compatible) backends.

ainter_*.yaml
model:
"llama3.2-internal":
"0.1:stable":
provider: "ollama"
model: "llama3.2"
format: "json"
base_url: "http://172.30.128.1:11434"
temperature: 0.0
openrouterautofree:
"0.1:stable":
provider: "openai"
openai_api_base: "https://openrouter.ai/api/v1"
openai_api_key: "$key:openrouter_openai_api_key"
model_name: "openrouter/free"
temperature: 0.0

System Component​

The System component defines individual context templates used to construct the comprehensive SYSTEM prompt context window. The layout example below tracks multiple modular namespaces optimized for specialized prompt assembly blocks. All examples are initialized at version 0.1 and aliased to the stable environment stage.

The system dictionary architecture is built to align with standard conversational frameworks (such as defining system persona context, background enterprise knowledge bases, strict structural constraints, chain-of-thought directives, etc.). However, following this explicit taxonomy is optional. The only structural requirement enforced by the compiler parser is that the third nesting layer must declare the block Name, and the fourth nesting layer must declare its Version.

ainter_*.yaml
system:
"persona": # Defines the operational role and behavioral constraints of the AI
"identify_speaker":
"0.1:stable": |
You are a Call Center call analysis specialist.
Your task is to analyze a call transcript and classify EACH detected 'SPEAKER_XXXX' EXCLUSIVELY into one of the following profiles:
1. AGENT: The human agent providing the service.
2. CUSTOMER: The person seeking support.
3. RECORDED_VOICE: Automated voices, menu options, or recorded messages.
4. IVR: Sub-group of RECORDED_VOICE oriented towards interactive menu choices, if differentiable.
5. ANNOUNCEMENT: Sub-group of RECORDED_VOICE oriented towards generic recorded announcements, if differentiable.
6. MUSIC: Segments identified as hold music, including background tracks with lyrics that are non-actionable for customer service.
7. AGENT2: A second (or subsequent) agent (e.g., supervisors, tier-2 support, or warm transfers).
8. CUSTOMER2: A second (or subsequent) person alongside the customer (e.g., a spouse, translator, or relative).
If the text content or context permits, please also extract and identify the specific names of the Customers and Agents.

"knowledge": # Static reference data and context injected into the context window
"explain_tags":
"0.1:stable": |
------------------------------------------
The call metadata, speaker index, and individual dialogue phrases may contain tags. In addition to the raw transcript text, these tags provide granular behavioral and environmental context for each speaker and phrase.
The available tags are:
1. mainspeaker: The primary speaker who dominated the call duration.
2. mainlanguage: The dominant language detected for the overall call or specific speaker.
4. mainlanguagevalue: Percentage score indicating the speaker's adherence to their primary language.
5. RECORDED_VOICE: Flag indicating the speaker is suspected of being an automated system based on voice stability and amplitude consistency (heuristic-based, may contain false positives).
6. MUSIC: Flag indicating the speaker segment is suspected of being background music based on repetitive wave patterns or rhythmic phrasing (heuristic-based, may contain false positives).
7. IVR: Flag indicating the speaker is suspected of being an IVR script based on typical conversational routing vocabulary (heuristic-based, may contain false positives).
8. ANNOUNCEMENT: Flag indicating the speaker is suspected of being a pre-recorded broadcast or advertisement based on static lexical cues (heuristic-based, may contain false positives).
9. UNCERTAIN_AUDIO: Highlights a high level of transcription ambiguity. Treat this segment with caution as it carries a risk of model hallucination.
10. BACKGROUND_NOISE: Indicates the phrase segment is suspected of containing high ambient or environmental background noise.
11. WEAK_SIGNAL: Indicates the phrase was captured at an extremely low volume, implying potential uncertainty regarding its operational relevance to the call.
12. AUDIO_CUTS: Indicates suspected digital audio clipping or packet drops, which may degrade readability.
13. DISTORTION: Flags severe harmonic distortion or audio saturation. This may impact transcription accuracy.
14. UNRECOGNIZED: Outlines almost imperceptible audio where the transcription engine carries maximum uncertainty.
15. DTMF: Identifies the phrase segment as suspected telephone keypad touch-tones.
16. LONG_SILENCE: Captures the occurrence of an unusually extended silence duration prior to the current phrase.
17. HIGH_TONE: Flags a phrase spoken at a pitch significantly higher than the speaker's historical baseline.
18. FAST_SPEAK: Flags a phrase delivered at an unusually high words-per-second velocity.
19. WHISPER: Flags a phrase spoken at an amplitude or pitch significantly lower than the speaker's historical baseline.
20. OVERLAY: Flags a phrase that physically overlaps with another active speaker channel. This may degrade transcription veracity or indicate cross-channel bleeding.
Tags flagged as suspicious should not be treated as absolute classifications, but rather as supportive structural markers to assist in profile identification.

"explain_userdata":
"0.1:stable": |
-------------------------------------------
The customer interaction dataset (passed downstream into the user prompt) is structured into 3 distinct components:
1. Call Header: High-level session data and operational metadata.
2. Speaker Directory: An indexed inventory of all detected speakers (SPEAKER_XXXX) along with their individual performance metrics.
3. Transcript Stream: A chronological sequence of dialogue phrases containing speaker maps, contextual tags, and the raw text payload.

"rules": # Core constraints defining what the model must or must not do during the interaction
"guardrails": # Explicit security boundaries, formatting rules, and alignment instructions
"chainofthought": # Methodological instructions forcing the AI to reason sequentially before responding
"ragfacts": # External context and dynamic knowledge retrieved from vector or relational databases
"examples": # Few-shot input/output interaction primitives to align model expectations
"outputformat": # Strict structural directives governing the response payload
"identify_speaker":
"0.1:stable": |
-------------------------------------------
Respond EXCLUSIVELY in a valid, structured JSON format matching the schema layout below. (Note: This is a static example of a call containing 3 active speakers; your actual payload must adapt dynamically to the exact number of speakers discovered in the source phrases):
{{
"speaker": {{
"SPEAKER_0000": {{"profile": "IVR", "justification": "Greets the user with a default automated welcome menu."}},
"SPEAKER_0001": {{"profile": "AGENT", "name": "John", "justification": "Explicitly identifies themselves as a support representative."}},
"SPEAKER_0003": {{"profile": "CUSTOMER", "justification": "Presents the core technical issue and account details."}}
}}
}}
Generate a clean JSON object containing only the speakers actively mapped within the dialogue phrases. Do not append any markdown wrappers, conversational filler, or trailing explanations outside of the strict JSON structure.

"structuralanchors": # Delimiters and tokens used to cleanly separate prompt context sections
"usercontext": # Dynamic attributes and persistent profile history regarding the current customer
"fallback": # Execution directives and default states when the AI cannot resolve an answer
"connector": # Interface bindings and pipeline integration tools
"tools": # Declarations for external tool-calling execution loops and third-party APIs
ainter_*.yaml
# Starting from version 0.2 of Ainter you can use Structured Outputs / JSON Schema for the LLM who support it (almost every one these days)
# Let's start with our example:

system:
"outputformat": # how to format the response
"identify_speaker":
"0.2:test":
speaker: # Root elemento
type: object # speaker is an object (dict)
description: "for every single speaker. the key is the speaker"
additionalproperties: # Use assitionalproperties to define the content of an object with a dynamin key
type: object # every speaker have an object
description: "information about a speaker"
properties: # Use properties to defibe the keys of an object
profile: # first key has the following content
type: string # profile is a string
description: "the profile of the speaker"
justification: # second key has the following comtent
type: string # justification is a string
description: "the justification for the choosen profile"
# type: object, array, str, int, float, bool
# description: for all types
# properties: defines the content pg every key elements of an object
# additionalproperties: defines the content of every dynamic keys element of an object
# item: define the content for every elemento of an array

User Component​

The User Component defines the modular blocks used to construct the USER prompt context window. The configuration structure matches the system component rules: the third nesting layer represents the component Name and the fourth layer declares its Version.

While you have complete flexibility over this structure, the specific identifiers speaker and dialog carry dedicated programmatic meanings within the engine. Inside an execution action, the runtime compiler evaluates loops to dynamically iterate and repeat data nested under speaker (for each designated call participant) and dialog (for each transcript phrase line). Therefore, using speaker and dialog as root keywords is mandatory if you intend to dynamically map multi-speaker metadata and interaction streams.

Furthermore, you can pull real-time data attributes directly from the source CTT file. Any expression wrapped inside { ... } tokens is processed dynamically by the Cequation engine. Within each name/version namespace, you can declare a single text payload string or an array of structural assignments.

When an array block is evaluated, the left side of the assignment determines whether the computed result will be outputted directly into the final prompt stream or temporarily cached inside a local runtime variable (such as using dialog and dialoghastags to store intermediate Cequation evaluations). Prefixed keys that explicitly start with a $ token (such as $print) are automatically rendered and outputted directly into the raw user prompt text payload.

To reference and evaluate an active local temporary variable within downstream expressions, you must prefix its namespace with a leading dot (e.g., .dialog). These fundamental processing rules apply symmetrically to system prompts, although runtime data mapping and variable injection are considerably more common within the user prompt context window.

ainter_*.yaml
user:
"call":
"metadata":
"0.1:stable": "Call Center call {$uuid} with {$origin.durationorig} seg. and with the following Tags: [{$call.tags::''.:join:''}]"
"speaker":
"metadata":
"0.1:stable": "Extract profiles exclusively for speakers with active dialogue entries. Omit any speakers with zero phrases. The active speakers in this interaction are detailed below by phrase count, duration (seconds), and metadata tags:"
"list":
"0.1:stable": "{$speaker.[my].$key} : {$speaker.[my].turns} / {$speaker.[my].duration}seg. [{$speaker.[my].tags::''.:join:''}]"
"dialog":
"metadata":
"0.1:stable":
- dialog: "[{$dialog.[my].tags::.:join:}]"
- dialoghastags: "{.dialog $dialog.[my].tags::.:join: '' != AND}"
- $print: "Phrase {$dialog.[my].$key} {$dialog.[my].speaker::UNKOWN} {.dialoghastags}"
"detail":
"0.1:stable": "{$dialog.[my].text}"
  • call.metadata – Fetches high-level transactional and structural metadata regarding the target call session.
  • speaker.metadata – Generates a semantic header descriptor introducing the speaker mapping matrix.
  • speaker.list – Loops through and appends explicit performance metrics for every speaker index registered by the execution action.
  • dialog.metadata – Dynamically assembles a structured context header for each dialogue line (repeated sequentially per entry).
  • dialog.detail – Injects the raw speech-to-text text payload associated with the corresponding dialogue phrase.

Output Component​

The Output Component defines the validation rules and structural paths used to parse the LLM's structured JSON response and patch the results cleanly into a newly compiled version of the CTT file. The configuration layout must target one or more baseline structural roots. A baseline target can map a single aggregation payload of global call values, evaluate an array block broken down per speaker, or process a sequential timeline tracking individual dialogue phrase blocks.

The configuration example below outlines a speaker-level parsing loop targeting two discrete key-value responses for each participant channel.

ainter_*.yaml
output:
"speakerret":
"0.1:stable":
"_speaker":
"?speaker.[myspeaker]":
- "$speaker.[myspeaker].profile": "{#output.speaker.[myspeaker].profile}"
- "$speaker.[myspeaker].justification": "{$output.speaker.[myspeaker].justification::}"

The _speaker namespace (and any command key prefixed with a leading underscore) functions exclusively as a structural code receptacle to split complex script nesting blocks into readable sub-sections. The statement ?speaker.[myspeaker] acts as an execution loop instruction, forcing the compiler to run the underlying assignments against every speaker reference using myspeaker as a dynamic iterator variable.

The two downstream assignment operations dynamically inject a profile and a justification attribute to the respective speaker indices within the target CTT ledger. The right side of the statement maps the incoming JSON payload extracted from the AI response, where the $output object references the root JSON string parsed from the LLM endpoint using the myspeaker variable.

Note on Loop Behavior: This baseline loop condition will attempt to cross-reference and evaluate every single speaker profile registered within the native CTT dataset. If an active speaker ID in the CTT file is missing from the incoming JSON object returned by the model, the parser will throw a validation error. If you prefer to exclusively iterate and process only the subset of speaker entities actively returned within the model payload, substitute the loop directive with ?output.speaker.[myspeaker] instead.


Action Component​

The Action Component acts as the top-level orchestration layer that binds specific system prompts, user prompt segments, model backends, and output schemas together into a unified executable pipeline task. When invoking the Ainter utility, you pass the explicit identifier of the target action you wish to deploy.

If you pass only the naked name of the action namespace (e.g., identify), the engine automatically falls back to loading the highest-indexed version bound to a stable tag alias. Alternatively, you can pin execution to an absolute version number (e.g., identify:0.1) or trigger targeted pipeline stages by mapping environment labels (e.g., identify:test).

ainter_*.yaml
action:
"identify":
"0.1:stable":
- "m.llama3.2-internal:stable"
- "s.persona.identify_speaker:stable"
- "s.knowledge.explain_tags:stable"
- "s.knowledge.explain_userdata:stable"
- "s.outputformat.identify_speaker:stable"
- "u.call.metadata:stable"
- "u.speaker.metadata:stable"
- "u.speaker[].list:stable"
- "u.dialog[]":
- ".metadata:stable"
- ".detail:stable"
- "{glue.: }"
- "o.speakerret:stable"

The Action builder systematically constructs and wraps the context window using explicit component prefixes:

  • m. – Maps the baseline Model gateway configuration to deploy (e.g., llama3.2-internal).
  • o. – Maps the targeted Output schema parser used for ingestion.
  • s. – Concatenates an ordered array of System blocks to compile the final SYSTEM prompt.
  • u. – Concatenates an ordered array of User blocks to compile the final USER prompt.

You must declare the absolute target path when stitching namespaces together. This can be achieved via single explicit strings (e.g., u.speaker.metadata) or assembled using relative block subdivisions (e.g., nesting .metadata under a parent u.dialog declaration). You can break out of a nested relative block and restart evaluation directly from the global root schema anywhere by re-declaring an absolute u. or s. prefix token.

The speaker and dialog keys allow you to pass custom array index boundaries []. Declaring an index token forces the runtime manager to replicate that prompt element according to specified indexing constraints:

  • [] – Evaluates all available speaker or dialogue entries in their natural chronological sequence (standard out-of-the-box pipeline behavior).
  • [variable:low_index:high_index:step] – Sorts the target dataset by the designated variable attribute and loops execution from the low_index boundary up to (but excluding) the high_index step delta.
    • Examples:
      • speaker[duration::-3:-1] – Sorts the dataset by length and triggers two loop repetitions, starting from the longest-duration speaker down to the second-longest speaker segment.
      • dialog[:0:2] – Restricts prompt rendering exclusively to the first two chronological dialogue phrase segments.

The syntax token {glue.: } merges the prompt output strings generated by the two immediately preceding instructions into a single unified sentence row, using the characters defined inside the expression (in this instance, : ) as the formatting delimiter.