-
Type:
New Feature
-
Resolution: Fixed
-
Affects Version/s: None
-
Component/s: translate5 AI
-
Medium
-
None
-
New feature where AI will suggest translation issues in the segments and with the help of MQM tags those issues are marked visible in editor.
-
None
-
Emptyshow more show less
Problem
Today's TQE (the OpenAI plugin) asks the AI, in one prompt, to estimate the quality of a batch of segments. The AI returns a numeric score plus free-text reasoning; translate5 stores the score in LEK_segments.qualityScore and in the TQE analysis results. This has two problems:
- The project manager cannot influence the score. The AI alone decides both what is wrong and how much each problem deducts. A PM cannot say "a terminology error costs 50%, a style error costs 10%".
- It is not transparent. You only get a number and prose reasoning - you cannot see which error, in which place, caused which deduction, nor act on it.
translate5 already has a complete MQM apparatus built for human linguists (an MQM error-type catalog, severities, inline MQM tags in the target text, per-segment quality rows, and a statistics view), but there is no weighted score computed from MQM annotations, and no way for the AI to produce MQM annotations. The MQM statistics today only count errors; they do not score them.
We want a second TQE mode where the AI does only the linguistic marking - identify translation errors and mark them as MQM issues - and translate5 owns the scoring with PM-defined weights. The result is transparent (every marked error is visible as MQM tags in the t5 editor), correctable (human-in-the-loop), and configurable.
Solution
TQE mode switch (AI | AI MQM). Reuse the existing TQE task assignment and the AI resource already assigned to the task - the same resource serves AI or AI MQM TQE. The mode is the config runtimeOptions.plugins.OpenAI.tqeMode (classic | mqm), overridable in the normal hierarchy at system, client and import level (not task level - once a task's score is computed, the calculation must not change underneath it). Default = classic (backward compatible).
MQM-tagging step (AI). In MQM mode the TQE worker/prompt tells the assigned AI resource the active MQM catalog and severities for the task and asks it to return, per segment, the within-segment MQM issues it finds - each marking a specific range inside the target text, not a whole-segment quality flag. Per issue: issue type, severity, the quoted target range, a confidence/probability, and a short reason. translate5 anchors each quoted range in the target text - issues whose quote cannot be located unambiguously are strictly discarded and logged, never guessed into a position. The accepted issues are mapped into the existing sub-segment MQM tags (the paired qmflag tags placed over the marked range) and persisted as normal within-segment quality rows with start/end positions, so reviewers see and can correct them in the editor (human-in-the-loop), exactly like manually placed MQM tags. Note: this is the within-segment (sub-segment) MQM, not translate5's whole-segment quality flags.
Scoring step (conventional, no AI). translate5 calculates the score from the MQM issues using the official MQM scorecard model (themqm.org scoring Excel): per segment, score = max(0, (1 - penalty points / word count) * 100), where every counted issue costs severity multiplier x error type weight penalty points and the word count is the segment's source word count. The AI's probability per issue does NOT scale the penalty - it only decides whether an issue counts at all (see probability threshold below); a counted issue always costs its full penalty, so the same issues always produce the same score. The resulting score is stored where classic TQE stores it (segment qualityScore plus reasoning summary plus the analysis results), so existing segment filters, KPI and statistics keep working unchanged.
Scoring-weights configuration. Extend the existing MQM criteria XML (QM_Subsegment_Issues.xml / QM_Subsegment_Issues_core.xml, imported per task into LEK_task.qmSubsegmentFlags) with a weight attribute per issue type (default 1, inherited from parent to child) and per-severity penalty multipliers as root attributes (e.g. severity-critical="25"). The whole XML is maintainable as the new config runtimeOptions.editor.qmFlagXmlContent - a config value holding the XML, edited in a validating XML editor window in the configuration UI and overridable at system, client and import level - rather than introducing a separate Excel/scorecard file.
Probability threshold configuration. The same XML additionally carries a probability threshold attribute per issue type (percent). AI issues whose probability is below their issue type's threshold are discarded at parse time; issue types without a threshold always count. The probability of counted issues is stored per quality row and shown in the score reasoning for transparency, but it never factors into the score itself.
Triggering. The same worker path can be started three ways:
- Automatically at import time (as today), mode-aware.
- A workflow action (QueueTqeAction) that queues the worker at a configured workflow event (e.g. after import/pre-translation), mirroring the existing Hotfolder ExportReadyProjectAction pattern. It ships opt-in: no default LEK_workflow_action row is installed - inserting a row per installation (choosing workflow and trigger there) is the activation. The user-facing "Automation Manager" UI is a separate, future feature.
- The existing manual "Start TQE" button, made mode-aware: when AI MQM TQE is configured for the task, clicking it re-runs the AI MQM tagging AND re-runs the score calculation. This lets a user re-check text quality at any point in the ongoing process (e.g. after post-editing, after review).
Re-running and calc-only. Presence of an AI/LLM resource decides what runs: AI resource selected = the AI re-tags the segments and the score is recalculated - this wipes ALL existing MQM tags of the task, including manually placed ones, before re-tagging. AI resource de-selected = only the calculation is re-run on the current MQM tags; manual corrections and false-positive marks survive. So if MQM tags were altered manually and the user just wants the score refreshed, they run "Start TQE" with the AI resource de-selected. The same rule applies to the workflow action: no AI resource means calculation only (and in classic mode without a resource, nothing runs).
Score refresh on segment save. When TQE on segment save is enabled (runtimeOptions.plugins.OpenAI.qualityScoreSegmentCheck) and the task is in MQM mode, every segment save recalculates the segment's score from its current MQM tags - a pure calculation without any AI call, regardless of whether an AI resource is assigned - and the editor grid receives the fresh score directly with the save response. The AI never re-tags on segment save; re-tagging only happens via the explicit triggers above.
Important implementation informations
- the current implementation follows the Raw Quality Score model from the MQM-2.0-Dual-Scorecard-2025-05-09 score card
Confluence documentation: https://confluence.translate5.net/spaces/BUS/pages/773947393/MQM-based+Translation+Quality+Estimation+TQE