-
Type:
Improvement
-
Resolution: Unresolved
-
None
-
Affects Version/s: None
-
Component/s: TermTagger integration
-
Medium
-
None
-
None
-
Emptyshow more show less
Problem
It should be possible to determine the exact spelling of a term in regard of lower/upper case.
Solution
- find and tag all terms regardless of upper/lower case, e.g. de "Sensor" and en "sensor" should tag text with de "SENSOR" and no term error given for usage of en "sensor" or usage of en "SENSOR".
- however, if in text exact match for lower/upper case is found, here "Sensor", then exact target term "sensor" is expected, else an error is thrown (for en "SENSOR" e.g.)
| term entry de | tern entry en | term in source text | term in target text | term error |
|---|---|---|---|---|
| Sensor | sensor | SENSOR | SENSOR | no |
| SENSOR | SENSOR | Sensor | sensor | no |
| SENSOR | SENSOR | SENSOR | SENSOR | no |
| SENSOR | SENSOR | SENSOR | sensor | yes |
| SenSor | SenSor | SenSor | sensor | yes |
Rule: exact casing match for term entry and source text -> exact target term casing required, else error
Scenario: If for SENSOR in source text capital letters in target are required, users need to maintain 2 term entries, one SENSOR - SENSOR and one Sensor - sensor
Exception: capitalization for term pretranslation for term-only segments:
term entry Sensor - sensor will be pretranslated in term-only segment with Sensor - Sensor. Here exact match is found, within a phrase "Sensor" will throw an error but in term-only segment it should not.
- causes
-
TRANSLATE-5694 Add term tagging via stanza support
- Done
- relates to
-
TRANSLATE-5180 Switching termtagging to Spacy.io and Stanza and refactoring termTagging code inside translate5
- In Progress