-
Type:
New Feature
-
Resolution: Unresolved
-
None
-
Affects Version/s: None
-
Component/s: translate5 AI
-
High
-
None
-
None
-
Emptyshow more show less
Problem
PONS and AI language resources should be able to consume terminology and TM matches produced by translate5 as additional context for their requests. Today the only implementation of this idea is OpenAI-internal: the connector fetches terms inline per segment via the TermTagger TerminologyProvider in three separate call sites with two different config gates (OpenAI/LanguageResource/Connector.php:761-770 for single queries, the ChunkPolicy terms closure for batches, QualityEstimate/QualitySegmentFactory.php:71-96 for TQE). For TM matches no reusable mechanism exists at all.
If every connector implemented its own internal searches of termCollections and translation memories, that logic would be duplicated and connectors would be tightly coupled to specific storage systems. The PONS technical concept defines a shared context architecture as the prerequisite for sending translate5 terminology and TM matches to PONS.
Solution
Introduce a reusable context package under application/modules/editor/src/LanguageResource/Context/.
Capability interfaces (connector side)
- SupportsTerminologyContext - setter-based injection of a prepared TerminologyContext
- SupportsTmContext - setter-based injection of a prepared TmContext
Connectors declare the capabilities independently of each other; a resource without context support is completely unaffected.
Context structures and providers
- DTOs: GlossaryTerm (source, target), TmMatch (source, target, similarityScore), TerminologyContext and TmContext (per-segment keyed collections)
- TerminologyContextProvider - wraps the existing TermTagger TerminologyProvider (Plugins/TermTagger/TerminologyProvider.php:110) and the task termCollection resolution (Models/TermCollection/TermCollection.php:99); no new term-search logic
- TmContextProvider - builds TM matches per segment from the batch-result cache (Pretranslation/BatchResult.php:53); becomes fully functional for t5memory with the later t5memory batch tickets, the architecture and provider are delivered here
Provider behaviour (max matches, minimum similarity, on/off) is configured on the providers - uniform for all consuming connectors, not per connector.
Orchestrator
An orchestrator that:
- Determines which context types a resource supports (capability interfaces)
- Requests those contexts from the appropriate providers for the segments at hand
- Injects the prepared context into the connector
- Invokes the connector only after all required contexts are available
Hook points: the single-segment funnel editor_Services_Connector::_query() (Services/Connector.php:210) covers the interactive editor and the non-batch analysis path in one place; the batch paths are hooked in the batch worker/BatchQueryService segment loop (Plugins/MatchAnalysis/BatchWorker.php:92, src/LanguageResource/Pretranslation/Batch/BatchQueryService.php:161-190) and the legacy BatchTrait loop. So batch analysis and interactive editor requests both use the same architecture.
Semantics
- Empty context is valid: "provided but empty" (empty collections) is distinguished from "not provided" (no injection), so a connector can distinguish translate5-provided context from the case where its own connector-specific fallback resources (e.g. PONS-side glossaries) should be used.
- Context data is injected - a connector never retrieves termCollections or TM data internally.