The important architectural decision is that the AI does not need to control the conversation. It observes, understands and suggests; the Supporter decides.
The foundation for real-time video and audio communication between the client and the Supporter.
Explored for real-time transcription with speaker identification, producing the conversation context.
Processes the conversation context and produces assistance for the Supporter.
The Supporter remains the final human actor. No AI output becomes a consequential action on its own.
Understanding what has already been discussed.
Identifying recurring topics and concerns.
Helping the Supporter think about what may be happening in the conversation.
Suggesting questions that could help the Supporter explore an issue.
Helping Supporters work within predefined frameworks and their training.
Producing structured post-session summaries.
Identifying information that may warrant additional attention.
The client should not need to interact with the AI. The Supporter should not become dependent on it. The conversation should continue even if the AI is unavailable.
Understand the current conversation.
Avoid unnecessary interruptions.
Make it clear when something comes from AI.
Assist rather than command.
Operate within predefined responsibilities.
The Supporter decides how to act.
A working area for the prompt engineering behind the Supporter assistant — kept alongside the product rather than buried in a repository.
Current production and development prompts.
How conversation history is provided to the model.
The expected structure of AI responses.
Different approaches tested, and what each changed.
How AI outputs are judged for usefulness and accuracy.
Irrelevant suggestions, hallucinations, incorrect interpretations, excessive intervention, missed context.
Prompts and evaluation notes are being written up as the prototype stabilises. Failure cases are collected from the start — they are the most useful record the project can keep.