Skip to content

Apply the MO header charset before decoding metadata - #1375

Open
fhgffy wants to merge 1 commit into
python-babel:masterfrom
fhgffy:fix/mo-header-charset
Open

fhgffy wants to merge 1 commit into
python-babel:masterfrom
fhgffy:fix/mo-header-charset

Conversation

@fhgffy

@fhgffy fhgffy commented Oct 9, 2026

Copy link
Copy Markdown

read_mo parses the Content-Type header as bytes, but decodes the metadata message with the default UTF-8 charset before applying that header. A Latin-1 MO containing non-ASCII project or translator metadata raises UnicodeDecodeError at that point.

Apply the already parsed Content-Type through the existing MIME-header setter before decoding the metadata entry. Add roundtrips for UTF-8, Latin-1 and cp1252 catalogs with non-ASCII metadata, message IDs, contexts and translations.

Validation: the regression fails on the original source; all 6 MO tests and the complete suite (7,829 passed, 10 skipped, 2 xfailed) pass after the fix. Project Ruff and whitespace/debug hooks and sdist/wheel builds passed.

Specification: https://www.gnu.org/software/gettext/manual/html_node/MO-Files.html

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant