-
Type:
New Feature
-
Resolution: Unresolved
-
None
-
Affects Version/s: None
-
Component/s: MatchAnalysis & Pretranslation, t5memory
-
High
-
None
-
None
-
Emptyshow more show less
Problem
A PONS (or AI) batch worker cannot use translate5 TM matches unless t5memory has already completed its searches and populated the batch-result cache. With TRANSLATE-5678 alone, the t5memory batch runs as one more instance of the same worker class as all other batch workers (editor_Plugins_MatchAnalysis_BatchWorker, queued once per language resource) - and the worker dependency system orders workers by class name per task (Zf_worker_dependencies + the blocking query in ZfExtended_Models_Worker:320-338). Two workers of the same class cannot be ordered against each other, so the required sequence "t5memory first, context-aware resources afterwards" cannot be expressed today. The technical concept identifies this ordering as mandatory.
Solution
Dedicated t5memory batch worker
Introduce a dedicated worker class (e.g. editor_Plugins_MatchAnalysis_TmBatchWorker, subclass of the existing batch worker - same work() logic, distinct class name for dependency targeting). The batch worker queueing (import/analysis path and pivot path) splits the batch-capable resources: TM resources (t5memory) are queued as TmBatchWorker, everything else stays BatchWorker.
The worker is only queued when a t5memory resource is actually associated with the task - tasks without assigned TMs create no additional processing.
Worker dependencies
New Zf_worker_dependencies rows (migration):
- editor_Plugins_MatchAnalysis_BatchWorker depends on TmBatchWorker - the context-aware PONS/AI batch workers do not start before all t5memory batch workers of the task reached a terminal state
- editor_Plugins_MatchAnalysis_Worker (and thereby everything downstream: pre-translation, TQE, automated post-editing) additionally depends on TmBatchWorker, analog to its existing dependency on the batch worker
Retry and failure behaviour
- Idempotent runs: at worker start the existing cache rows for (task, language resource) are deleted before new results are written - retries can never produce inconsistent duplicate cache entries (the cache read already takes the latest row, this makes re-runs clean at the source).
- Failure: if the t5memory worker ends defect after its retries, that is a terminal state - dependent workers proceed and find an empty/partial cache, which is a valid empty TM context downstream. The failure itself is logged task-visible, so the "context ran without TM matches" situation is diagnosable. No hard block of the whole import chain because of a TM outage.