-
Type:
Bug
-
Resolution: Unresolved
-
None
-
Affects Version/s: None
-
Component/s: t5memory
-
High
-
None
-
Improve processing of invalid XML characters
-
None
-
Emptyshow more show less
Problem
In the W3C XML 1.0 specification, hexadecimal numeric character references must strictly use a lowercase x:
$$\text{CharRef} ::= \text{'&#' } [0-9]+ \text{ ';' } \mid \text{'&#x' } [0-9a-fA-F]+ \text{ ';'}$${code}
Some CAT tools are known to export XML/TMX files using an uppercase *{{&#X}}* (for example, {{*方*}} instead of {{{}*方*{}}}). While HTML5 tolerates uppercase {{{}*&#X*{}}}, XML parsers (such as Java's Xerces parser used by t5memory) strictly reject it as invalid XML with the exact fatal error:
{quote}{{Fatal Error: hex radix character references must use 'x', not 'X'}}{quote}
In test file, the 7 Chinese characters (e.g. {{{}方舟美服项目组{}}}, which appears in {{{}creationid{}}}) or other non-ASCII characters were written in the raw file as:
{code:java}
方舟美服项目组
Solution
Add processing for TMX import to replace &#X with &#x