-
Type:
New Feature
-
Resolution: Unresolved
-
None
-
Affects Version/s: None
-
Component/s: translate5 AI
-
Medium
-
None
-
None
-
Emptyshow more show less
Summary
Add a self-learning AI prompt capability to translate5 AI based language resources (both translation and TQE). The feature automatically proposes improved RAG sysMessages by learning from real post-edits done by users in finished tasks. Post-edits from trusted editors weigh more. Proposals appear in the RAG UI where users can promote/enable them.
User perspective
As a translate5 admin managing OpenAI language resources, We want translate5 to propose improved prompts based on real post-edits from trusted users, so that the resource's output gets progressively closer to what good linguists actually produce without writing prompts by hand.
Idea how it should work
1. Trust score for editors.
Each user gets a score reflecting how reliable their post-edits are as a learning signal - high score =
their post-edits carry more weight. The score is derived from existing post-editing aggregation data:
consistency of their edits, time-per-segment compared to peers, and completed-segment volume. New users
get a neutral default so they're not punished for lack of history. The score is recomputed
automatically; admins do not maintain it by hand.
2. Collecting the learning data.
For each translate5 AI language resource, translate5 collects pairs of (model output, user's post-edit) from
finished tasks that used the resource. Each pair carries the trust score of the editor who produced it.
This collected dataset is the "ground truth" the self-learning AI prompt tries to get closer to over
time.
3. The tuning loop (runs once per resource per trigger).
- Take the resource's current RAG sysMessage and the collected dataset.
- Hold back 20% of the data as a frozen test set the tuner does not see while proposing.
- Identify the cases where the current prompt produced output furthest from what trusted editors
actually wrote - these are the cases that most need improvement. - Ask an LLM to propose a small number of revised prompts targeting those weak spots.
- Score each proposed prompt against the held-out 20%, weighted by editor trust.
- If a candidate clearly beats the current prompt on the held-out set, save it as a proposed
self-learning AI prompt for that resource. No automatic activation.
4. How the tuning is triggered.
Two trigger paths, both queueing the same background worker:
- Manual: user clicks a button on the language resource to tune it now.
- Automatic daily: a daily action checks every RAG-based translate5 AI resource and tunes
the ones that have accumulated enough new post-edits since their last run.
5. Human acceptance - the tuner never auto-activates a prompt.
A proposed self-learning AI prompt is shown to the user alongside the currently active RAG sysMessage
in the existing RagBased UI, with a side-by-side diff and the improvement metrics (how much
closer to post-edits it gets, on how many samples, how it was triggered). The user decides:
- Promote -> the proposed prompt becomes the active one for the resource. The previous one is archived,
not deleted. - Reject -> the proposal is discarded. The active prompt stays untouched.