-
Type:
Improvement
-
Resolution: Unresolved
-
None
-
Affects Version/s: None
-
Component/s: translate5 AI
-
High
-
None
-
Enhancement: Add automatic post-editing worker to solve problems detected in the TQE step
-
None
-
Emptyshow more show less
Problem
Newer AI models tend to answer meta-comments like "the translation is ambiguous because this and that" or "The segment "Sth." could not be translated because bla bla" instead of actually translating a segment.
Solution
1. Config is: Pretranslation by LLM and TQE by LLM on task import
Detect segments with
- TQE score below 50 and
- target text length is significantly longer
- look for hallucination-specific words in target (e.g. "translation")
Add an automatic post-editing worker after the TQE step to re-translate the segments where such problems (and other model hallucinations) are detected (e.g. by detecting an unnatural text-length together with the TQE score and the reasoning). The worker should retry to translate the segments with additional instructions. If 2 LLMs were used, either LLM 1 or LLM 2 can be used to re-translate.
2. Config is: Pretranslation by DeepL and TQE by LLM on task import
Detect segments with
- TQE score below 50 and
- target text length is significantly longer
- look for hallucination-specific words in target (e.g. "translation")
Add an automatic post-editing worker after the TQE step to re-translate the segments where such problems (and other model hallucinations) are detected (e.g. by detecting an unnatural text-length together with the TQE score and the reasoning). The worker should retry to translate the segments with additional instructions.The TQE LLM resource should be used to re-translate.
3. Config is: Pretranslation by LLM 1
Detect segments with
- target text length is significantly longer
- look for hallucination-specific words in target (e.g. "translation")
Add an automatic post-editing worker after the pre-translation step to re-translate the segments where such problems (and other model hallucinations) are detected (e.g. by detecting an unnatural text-length together with the TQE score and the reasoning). The worker should retry to translate the segments with additional instructions. Send it to the pretranslation LLM 1 with our TQE prompt.
For all 3
If the TQE value returned is below 50, resend it to the LLM for translation with the information of the TQE reasoning added and ask it to really send back a proper translation. And do this in loop 3 times, if answer still x% too long. And if then still nothing valid returned, log a warning and leave the segment empty for normal task and leave the hallucinations for InstantTranslate. If difficult to distinct this, leave the hallucinations everywhere.
- relates to
-
TRANSLATE-5280 Automated post-editing
- In Progress