-
Type:
Improvement
-
Resolution: Unresolved
-
None
-
Affects Version/s: None
-
Component/s: TermTagger integration
-
High
-
None
-
Improve case sensitive term handling
-
None
-
Emptyshow more show less
Problem
Terms in Term collection may have different cases and be written in different way.
Some of them should be handled differently regarding case and we don't have a good way for now of how to setup this behaviour in our App.
Solution
Different terms need different capitalization rules. An acronym such as IT should not necessarily match the pronoun it, while a common word should match sentence-initial and uppercase variants. Users need to control this separately for each term, including synonyms and source/target terms in the same entry.
Introduce the term-level attribute caseSensitivity, with serialized values soft, strict, and no. Display the labels Soft, Strict, and No under Case sensitivity. Use “Soft” consistently; the task’s final reference to “Permissive” describes the same mode and should be renamed.
- P1 Conservative case alignment for inflections/fuzzy edits, including lowercase suffixes, deleted capitals, inserted uppercase letters, and ambiguous alignments.
- P2 Canonical normalization plus locale-independent full Unicode folding, without automatic Turkish tailoring. The contract includes approved expected outcomes.
Matching contract
Case matching is directional: compare the stored term against the actual occurrence in the segment.
| Mode | Rule for otherwise identical spellings | Examples |
| — | — | — |
| Soft | Uppercase letters in the stored term must remain uppercase at their corresponding positions. Lowercase letters may appear in either case. | `Editor-in-chief` matches `Editor-in-chief`, `Editor-in-Chief`, `EDITOR-IN-CHIEF`; rejects `editor-in-chief`. `it` matches `it`, `It`, `IT`. |
| Strict | Letter case must be identical; this does not disable stemming or fuzzy matching. | Among case-only variants, `TermTagger` matches only `TermTagger`; `IT` rejects `it`; `Can` rejects `CAN` and `can`. |
| No | Ignore differences in letter case. | `TermTagger` matches `TermTagger`, `termtagger`, `TERMTAGGER`, `TermTagGer`. |
The policy does not itself relax punctuation, diacritics, whitespace, or word boundaries. Existing tokenization remains responsible for finding candidate spans. Compare using Unicode-aware rules and preserve offsets into the original segment. Do not implement this with ASCII-only lowercase conversion or modify the segment text.
Stemming and fuzzy matching remain controlled by their existing settings in all three modes. Case sensitivity is an additional constraint on those candidates, not an exact-spelling switch. In particular, Strict must not reject a match merely because the accepted inflected/fuzzy form has a different spelling or length. Illustrative acceptance case: if the existing stemmer returns `books` for `book`, Strict accepts the lowercase occurrence but rejects `Books` and `BOOKS`.
For inflections, substitutions, insertions/deletions, and mixed-case terms, the matcher must retain or expose enough correspondence between the stored term and the occurrence to evaluate capitalization. A position-by-position comparison of raw strings is insufficient. The approved P1 contract now specifies aligned letters, added lowercase letters, rejected added uppercase letters in Strict, deleted required capitals, and disputed alignment outcomes, including mixed-case/acronym suffixes. Do not use exact string equality or only a first-letter comparison as a shortcut.
Approved P2 Unicode baseline: canonical normalization and locale-independent full Unicode folding for comparison copies, with original case/offset provenance. The contract matrix specifies Turkish dotted/dotless I, German sharp S, Greek sigma, titlecase characters, and composed/decomposed accents. These comparisons do not broaden a provider's lexical candidate generation. Scripts without letter case should behave identically in all modes for otherwise identical spellings.
Apply the policy independently to source and target terms. A wrong-case target occurrence rejected by its policy must not satisfy terminology QA for a valid source occurrence. A source occurrence rejected by its policy must not create a required translation. Preserve existing preferred/forbidden status handling for accepted candidates.
- relates to
-
TRANSLATE-5694 Add term tagging via stanza support
- Done