Problem
test_html_format (openlibrary/i18n/test_po_files.py, and the copy in openlibrary-i18n tests/) compares a msgid's and msgstr's HTML trees tag by tag, in position. Languages whose word order differs from English legitimately reorder elements, and the test rejects those correct translations.
Found on #13704 (Azerbaijani): two correct translations reorder <b>/<em> and two <a> elements because Azerbaijani is verb-final. #13748 had to clear them to turn master green. In openlibrary-i18n, the pipeline's own gate keeps dropping the same class of string.
The better the translation, the more English we ship.
Decision needed
Relax the comparison to tag names, attributes and nesting as a multiset rather than by position? Or keep positional matching and accept the loss? Applies to both repos.
Problem
test_html_format(openlibrary/i18n/test_po_files.py, and the copy in openlibrary-i18ntests/) compares a msgid's and msgstr's HTML trees tag by tag, in position. Languages whose word order differs from English legitimately reorder elements, and the test rejects those correct translations.Found on #13704 (Azerbaijani): two correct translations reorder
<b>/<em>and two<a>elements because Azerbaijani is verb-final. #13748 had to clear them to turn master green. In openlibrary-i18n, the pipeline's own gate keeps dropping the same class of string.The better the translation, the more English we ship.
Decision needed
Relax the comparison to tag names, attributes and nesting as a multiset rather than by position? Or keep positional matching and accept the loss? Applies to both repos.