-
Type:
Improvement
-
Resolution: Fixed
-
Affects Version/s: None
-
Component/s: TermTagger integration
-
High
-
None
-
Redo Term Tagger integration logic.
-
None
-
Emptyshow more show less
Problem
For now biggest part of term tagging logic is done by TermTagger itself. But term tagger has known bugs.
Solution
Because of that we though to redo logic on translate5 side and use TermTagger as stupid word finder engine and do term matching on our own.
For this we send TBX to TermTagger as planar list of term without translation.
Each term will have its own entry node.
Homonyms will be represented as 1 term in TBX, list of ids from DB will be set as term id in TBX as concatenated list "|" separated.
Then on translate5 side we will compare ids in source and target for each found word-term and find pairs.
This will ensure that pair will be made only from terms of same entry.
Make sure that if term has capitalisation in source - in target term with capitalisation will also be chosen as preferable if exists.
As part of implementation we have to make preparations for system to be able to work with other engines that can find terms except of TermTagger.
- blocks
-
TRANSLATE-5180 Switching termtagging to Spacy.io and Stanza and refactoring termTagging code inside translate5
- In Progress
- relates to
-
TRANSLATE-5748 AI terminology lookup fails for edited segments (duplicate "target" field)
- Done