t5memory: Support batch analysis and cache TM matches per segment

XMLWordPrintable

    • Type: New Feature
    • Resolution: Unresolved
    • None
    • Affects Version/s: None
    • Component/s: t5memory

      Problem

      PONS and other context-aware resources require TM matches to be available before their own batch requests are created. t5memory currently does not participate in the batch-processing sequence at all: during analysis it is queried live per segment, and it never writes the batch-result cache - so there is nothing a TmContextProvider could read and no defined point in the worker sequence where t5memory results are complete.

      Technically t5memory cannot be batch-capable today because the legacy batch infrastructure treats a batch size of one as "not batch-capable" (BatchTrait::isBatchQuery() returns batchQueryBuffer > 1), and t5memory does not support multi-segment searches yet.

      Solution

      Enable t5memory to use the (legacy) batch connector infrastructure with a batch size of one, storing its results in the existing batch-result cache with an unambiguous association to the originating segment (segmentId + languageResourceId + taskGuid - the existing cache key). The proof of concept on t5memory-batch-result-cache validates the approach and is productionised in this ticket:

      • BatchTrait::isBatchQuery() accepts a batch size of one (>= 1 instead of > 1). Verified safe: all other trait users configure buffers of 20-100, so only t5memory changes behaviour.
      • The t5memory connector uses editor_Services_Connector_BatchTrait with batchQueryBuffer = 1; its batchSearch() delegates each entry to the existing query($segment) - match values, penalties and metadata are produced by exactly the same code as today and therefore remain unchanged.
      • Segment access in batchSearch(): instead of widening the batchSearch(string[] ...) contract to a multi-dimensional array for all connectors (the PoC's open TODO), the trait exposes the current batch entries via one protected property, so t5memory can read the segment objects and no existing batch connector has to be touched.
      • Results are written through the existing saveBatchResults() path into the batch-result cache; cached results are retrievable per segment (basis for TmContext).
      • An analysis run where t5memory finds nothing produces a valid empty cached result (empty context is valid downstream).
      • Ordinary interactive editor searches are unaffected: the cache path is only active in the batch-enabled analysis context; the editor keeps querying t5memory live.

      Note on purpose: with a batch size of one there is no speed gain - one request per segment as today. The point of this ticket is that t5memory participates in the batch worker sequence and populates the result cache, so that the dedicated worker ordering (follow-up ticket) can guarantee TM matches exist before context-aware resources run.

            Assignee:
            Aleksandar Mitrev
            Reporter:
            Aleksandar Mitrev
            Aleksandar Mitrev Aleksandar Mitrev
            None
            Axel Becher
            None
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

              Created:
              Updated:
              None
              None