translate5 AI: Add a "post TQE worker" / automatic post-editing to solve problems with meta-comments

XMLWordPrintable

    • Type: Improvement
    • Resolution: Unresolved
    • None
    • Affects Version/s: None
    • Component/s: translate5 AI
    • High
    • None
    • Enhancement: Add automatic post-editing worker to solve problems detected in the TQE step
    • None

      Problem

      Newer AI models tend to answer meta-comments like "the translation is ambiguous because this and that" or "The segment "Sth." could not be translated because bla bla" instead of actually translating a segment.

      Solution

      1. Config is: Pretranslation by LLM and TQE by LLM on task import

      Detect segments with

      • TQE score below 50 and
      • target text length is significantly longer
      • look for hallucination-specific words in target (e.g. "translation")

      Add an automatic post-editing worker after the TQE step to re-translate the segments where such problems (and other model hallucinations) are detected (e.g. by detecting an unnatural text-length together with the TQE score and the reasoning). The worker should retry to translate the segments with additional instructions. If 2 LLMs were used, either LLM 1 or LLM 2 can be used to re-translate.

      2. Config is: Pretranslation by DeepL and TQE by LLM on task import

      Detect segments with

      • TQE score below 50 and
      • target text length is significantly longer
      • look for hallucination-specific words in target (e.g. "translation")

      Add an automatic post-editing worker after the TQE step to re-translate the segments where such problems (and other model hallucinations) are detected (e.g. by detecting an unnatural text-length together with the TQE score and the reasoning). The worker should retry to translate the segments with additional instructions.The TQE LLM resource should be used to re-translate.

       

      3. Config is: Pretranslation by LLM 1

      Detect segments with

      • target text length is significantly longer
      • look for hallucination-specific words in target (e.g. "translation")

      Add an automatic post-editing worker after the pre-translation step to re-translate the segments where such problems (and other model hallucinations) are detected (e.g. by detecting an unnatural text-length together with the TQE score and the reasoning). The worker should retry to translate the segments with additional instructions. Send it to the pretranslation LLM 1 with our TQE prompt. 

       

      For all 3

      If the TQE value returned is below 50, resend it to the LLM for translation with the information of the TQE reasoning added and ask it to really send back a proper translation. And do this in loop 3 times, if answer still x% too long. And if then still nothing valid returned, log a warning and leave the segment empty for normal task and leave the hallucinations for InstantTranslate. If difficult to distinct this, leave the hallucinations everywhere.

       

       

       

       
       

            Assignee:
            Aleksandar Mitrev
            Reporter:
            Axel Becher
            Aleksandar Mitrev Aleksandar Mitrev
            Sylvia Schumacher
            Axel Becher
            Stephan Bergmann
            Votes:
            0 Vote for this issue
            Watchers:
            4 Start watching this issue

              Created:
              Updated:
              None
              None