Repository navigation
Fix requirements.txt with unmodelled coding line (#1212) - #1218
Mikola Lysenko (mikolalysenko) merged 3 commits into
Conversation
Assisted-by: Claude Code:claude-opus-5-5
A plain-ASCII requirements.txt whose PEP 263 header names a codec such as iso-8859-15, cp1250, mac-roman or gbk installs fine with pip, but socket-patch reads it as unreadable: scan finds no packages and exits 0, hosted get wires nothing, and vex and rollback fail on an existing hosted pin. These tests capture that through the decoder, VEX discovery, lock-only scan (root file and -r include, hosted and vendored) and the hosted get / vex / rollback round trip. They fail on main. Refs #1212 Assisted-by: Claude Code:claude-opus-5-5
A requirements.txt whose PEP 263 header names a codec such as iso-8859-15, cp1250, mac-roman or gbk was read as unreadable, even when the file itself is plain ASCII (a header copied from a template). Scan found no packages and exited 0 while pip installed the unpatched release; hosted scan and get wired nothing; vex and rollback could not see an existing hosted pin. pip decodes with the codec the line names, and every ASCII-compatible codec decodes ASCII bytes as ASCII. The decoder now knows every Python name for those codecs (the ISO-8859, Windows, DOS, Mac, KOI8 and CJK families), looked up the way Python normalizes a codec name, and reads plain-ASCII files under them. Non-ASCII bytes under those codecs, codecs that read ASCII differently (EBCDIC, UTF-16/32, UTF-7, HZ, ISO-2022, escape codecs) and names Python does not know still read as unreadable. Fixes #1212 Assisted-by: Claude Code:claude-opus-5-5
|
BugBot review Generated by Claude Code |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit f1f51da. Configure here.
|
Ready for review at
Generated by Claude Code |
Final review briefWhat it does: A requirements file whose PEP 263 header names a codec other than UTF-8, ASCII, Latin-1 or cp1252 (for example Risk: low. One new arm in Look here:
Verified:
Changes I made: none. Open questions: none. One non-blocking nit: Auto-merge (squash) is armed, so approving sends this straight to the merge queue. Generated by Claude Code |
LLM Description written by Claude Code:claude-opus-5-5
Fixes #1212
Summary
A requirements.txt whose PEP 263 header names a codec beyond UTF-8 / ASCII / Latin-1 / cp1252 (
iso-8859-15,cp1250,mac-roman,gbk, …) is read the way pip reads it again. Before this, a plain-ASCII file with such a header (copied from a template) was treated as unreadable. Scan found no packages and exited 0 while pip installed the unpatched release, hosted scan/get wired nothing, and vex/rollback failed on an existing hosted pin.Root cause
#1152 sent every requirements reader through
utils::requirements::decode. Itscoding_line_codeconly knew four codecs and returnedNonefor every other name, sodecodegave up on the whole file. pip'sauto_decodedecodes with whatever codec the line names, and every ASCII-compatible codec decodes ASCII bytes as ASCII. Every caller (the-rinclude walk behind lock-only discovery and the hosted rewriter, VEX discovery, rollback, the vendored probes) goes throughdecode, so this one fix covers all the symptoms.Change
coding_line_codecgains aCodec::AsciiCompatiblearm. It covers every Python 3.11 name (module and alias) for codecs that decode each ASCII byte as itself: the ISO-8859, Windows, DOS, Mac, KOI8 and CJK multibyte families, 288 names. Names are looked up the way Python'snormalize_encoding/search_functionspell them. I generated the table from Python's own codec registry and checked each codec against all 128 ASCII bytes. Under these codecs, plain-ASCII bytes decode as ASCII.latin-9example: Python's codec lookup rejectslatin-9(codecs.lookup('latin-9')raisesLookupError; onlylatin9/l9/iso-8859-15are aliases), so pip can't install from that file either. It stays unreadable, matching pip, and the tests uselatin9instead.Out of scope
A file with non-ASCII bytes under an unmodelled codec (e.g. a real cp1250 comment) is still unreadable. That is unchanged from before #1152 and needs per-codec tables; not needed for #1212's reported cases.
Test evidence
New tests, red on main (commit bc0d34c), green with the fix (f1f51da):
utils::requirements::tests::decode_reads_ascii_under_any_ascii_compatible_coding_line-rinclude, hosted + vendored)scan_requirements_lock_only::lock_only_scan_discovers_ascii_pins_under_any_ascii_coding_linelockfileOnlyPackages: 0)vex::discover::pypi_other::tests::requirements_files_under_an_ascii_compatible_coding_line_are_readmanifest_not_foundin_process_rollback_hosted::pypi_requirements_hosted_round_trip_with_an_unmodelled_coding_lineLocal runs:
cargo clippy --workspace --all-features -- -D warnings: clean.rustfmt --checkon the touched files: clean.cargo fmt --all -- --checkis already red onmainacross 17 untouched files, and CI doesn't run it.cargo test -p socket-patch-core --all-features --lib: 5889 passed. The other 4 failures are read-only-permission tests that can't fail when run as root in this sandbox (copy_treerelax loop,vlt_healunremovable lock, poetry/requirements wire-write failure). Untouched code.scan_requirements_lock_only,in_process_rollback_hosted,mode_migration_pypi,in_process_get_hosted_ecosystems,in_process_redirect_pipenv,vendor_eject_fresh_checkout, …): 194 passed. The 1 failure (pipenv_hosted_to_vendored_names_the_unpatched_requirements) needs a live GET to pypi.org, which this sandbox can't reach.cargo test --workspaceran out of sandbox disk while linking. CI runs the full matrix.No wrapper (
npm/,pypi/,gem/) changes needed: the decoder is Rust-only.🤖 Generated with Claude Code
https://claude.ai/code/session_01GwaJm5RXR9MjWKencyZifJ
Note
Low Risk
Narrow decoder fix for template coding headers on ASCII files; behavior for non-ASCII and exotic codecs is unchanged, with broad regression tests.
Overview
Fixes #1212: plain-ASCII
requirements.txtfiles whose PEP 263 header names a codec outside UTF-8 / ASCII / Latin-1 / cp1252 (e.g.iso-8859-15,cp1250) are decoded again instead of being treated as unreadable.utils::requirements::decodegainsCodec::AsciiCompatible, backed by a large alias table (ASCII_COMPATIBLE_CODECS) andpython_codec_keynormalization so names match Python’s codec lookup. For those codecs, ASCII-only bytes decode like pip; non-ASCII content and codecs that do not map ASCII 1:1 (EBCDIC, UTF-16, unknown names) stay unreadable.Regression coverage spans the decoder unit test, lock-only scan (root and
-rincludes), VEX discovery, and hosted get / vex / rollback round-trips.Reviewed by Cursor Bugbot for commit f1f51da. Configure here.
Generated by Claude Code