-
Type:
Bug
-
Resolution: Fixed
-
Affects Version/s: None
-
Component/s: Editor general
-
High
-
None
-
Fix for a problem where decoding protected content can wrongly evaluate non protected/placeable tags.
-
None
-
Emptyshow more show less
Problem
In a specific setup, querying an LLM/MT resource that uses the XLIFF tag dialect can return a segment where most of the translated text is missing, even though the resource itself answered correctly. The text goes missing on our side, while the internal tags are restored from the answer.
The setup: a segment contains a paired formatting tag (e.g. a bold bpt/ept pair) whose both halves carry the t5placeable CSS class. The XLF import writes this marker when runtimeOptions.import.xlf.placeablesXpathes is configured with a very broad XPath such as //* - the pair content <b> counts as placeable content because b is an allowed formatting tag. This exists since the placeables feature (TRANSLATE-3533) and never had an effect before, since the editor's placeable handling only considers single tags.
The protected-content tag-shape codec (TRANSLATE-5360) is the first code that reads the marker without the single-tag condition. It noted the pair ids in its EncodingContext although no wrapper was created for them on the wire (only singleton <x/> tags can be wrapped). When the answer comes back, the wrapper detector - which by design trusts the noted ids - takes the regular bx/ex pair for a protected-content wrapper and replaces the whole span with the original tag, so the translated text between the two formatting tags is not taken over into the result.
While investigating, a small unrelated issue was found in the OpenAI plugin: XliffReader::prepareXliff() had the str_starts_with() arguments swapped, so the leftover fence marker (xml/xliff) of a markdown-fenced LLM answer was never stripped (no visible effect so far).
Solution
One small, targeted change in the codec: XliffProtectedContentCodec::encode() now notes an id in the EncodingContext only when a wrapper was actually created for it on the wire. Ids without a wrapper can no longer reach the context, so the decode step simply leaves the formatting pair alone and the translated text is taken over as expected. This works for existing tasks as they are - no re-import, no config change, no migration.