PONS connector: Use matches from assigned t5memory resources

XMLWordPrintable

    • Type: New Feature
    • Resolution: Unresolved
    • None
    • Affects Version/s: None
    • Component/s: Editor general

      Problem

      When local t5memory matches exist for a segment, PONS should use them directly as contextual examples instead of querying a translation memory on the PONS side. The matches live in translate5; PONS supports inline translation_memory_matches in its translation requests exactly for this case.

      Important API behaviour: PONS treats inline translation_memory_matches and translation_memory_ids as mutually exclusive - when inline matches are present, the TM-ID vector search is skipped. Sending both in one request must therefore never happen.

      Solution

      The PONS connector implements SupportsTmContext (context architecture). The injected TmContext entries are converted per segment to:

      translation_memory_matches[]:
      - source_text
      - target_text
      - similarity_score

      Context sources (both already delivered by the architecture):

      • During task analysis/pre-translation, the context comes from the t5memory batch cache (populated and ordered by the t5memory batch worker tickets).
      • During editor searches, the context comes from the ordered interactive query pipeline (live t5memory queries inside the PONS request).

      Priority per request (using the guarded TM-context attachment spot in the request builder):

      1. If local t5memory matches exist, send translation_memory_matches.
      2. Otherwise, use the matching translation_memory_ids found by the automatic PONS resource selection.
      3. Otherwise, send neither field.

      Never send translation_memory_matches and translation_memory_ids together. For batch chunks this is enforced structurally: the chunk building partitions segments by context presence (segments with local matches vs. segments without), so every request is cleanly in one mode - segments without matches keep the PONS TM-ID fallback instead of losing it to a mixed chunk. Editor requests are single-segment and trivially clean.

      Mapping rules:

      • Similarity: translate5 matchrates above 100 (context/repetition matches 101-104) are clamped to 100; the PONS similarity scale (integer 0-100 vs. float 0-1) must be verified against the API and mapped consistently.
      • Content: match source/target are sent as plain text (internal tags stripped); whether PONS supports markup inside TM examples must be verified, until then plain text is the safe default.
      • Per-segment association: matches are attached to the segment they belong to (the exact request shape for per-segment matches must be verified against the PONS API).

      No complete TM is ever uploaded or synchronised to PONS.

            Assignee:
            Aleksandar Mitrev
            Reporter:
            Aleksandar Mitrev
            Aleksandar Mitrev Aleksandar Mitrev
            None
            Axel Becher
            None
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

              Created:
              Updated:
              None
              None