Skip to content

Fix requirements.txt with unmodelled coding line (#1212) - #1218

Merged
Mikola Lysenko (mikolalysenko) merged 3 commits into
mainfrom
agent/fix-requirements-unmodelled-coding-line
Oct 9, 2026
Merged

Mikola Lysenko (mikolalysenko) merged 3 commits into
mainfrom
agent/fix-requirements-unmodelled-coding-line

Conversation

@mikolalysenko

@mikolalysenko Mikola Lysenko (mikolalysenko) commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

LLM Description written by Claude Code:claude-opus-5-5

Fixes #1212

Summary

A requirements.txt whose PEP 263 header names a codec beyond UTF-8 / ASCII / Latin-1 / cp1252 (iso-8859-15, cp1250, mac-roman, gbk, …) is read the way pip reads it again. Before this, a plain-ASCII file with such a header (copied from a template) was treated as unreadable. Scan found no packages and exited 0 while pip installed the unpatched release, hosted scan/get wired nothing, and vex/rollback failed on an existing hosted pin.

Root cause

#1152 sent every requirements reader through utils::requirements::decode. Its coding_line_codec only knew four codecs and returned None for every other name, so decode gave up on the whole file. pip's auto_decode decodes with whatever codec the line names, and every ASCII-compatible codec decodes ASCII bytes as ASCII. Every caller (the -r include walk behind lock-only discovery and the hosted rewriter, VEX discovery, rollback, the vendored probes) goes through decode, so this one fix covers all the symptoms.

Change

  • coding_line_codec gains a Codec::AsciiCompatible arm. It covers every Python 3.11 name (module and alias) for codecs that decode each ASCII byte as itself: the ISO-8859, Windows, DOS, Mac, KOI8 and CJK multibyte families, 288 names. Names are looked up the way Python's normalize_encoding / search_function spell them. I generated the table from Python's own codec registry and checked each codec against all 128 ASCII bytes. Under these codecs, plain-ASCII bytes decode as ASCII.
  • These still read as unreadable, as before:
    • Non-ASCII bytes under those codecs (there is no table to decode them).
    • Codecs that read ASCII differently: EBCDIC, cp864, UTF-16/32, UTF-7, HZ, ISO-2022, Shift_JIS-2004, the escape codecs, idna and punycode.
    • Names Python doesn't know.
  • Note on the issue's latin-9 example: Python's codec lookup rejects latin-9 (codecs.lookup('latin-9') raises LookupError; only latin9 / l9 / iso-8859-15 are aliases), so pip can't install from that file either. It stays unreadable, matching pip, and the tests use latin9 instead.

Out of scope

A file with non-ASCII bytes under an unmodelled codec (e.g. a real cp1250 comment) is still unreadable. That is unchanged from before #1152 and needs per-codec tables; not needed for #1212's reported cases.

Test evidence

New tests, red on main (commit bc0d34c), green with the fix (f1f51da):

Issue symptom Test Red on main Green
decoder reads the file as absent utils::requirements::tests::decode_reads_ascii_under_any_ascii_compatible_coding_line FAILED ok
lock-only scan finds 0 packages, exit 0 (root + -r include, hosted + vendored) scan_requirements_lock_only::lock_only_scan_discovers_ascii_pins_under_any_ascii_coding_line FAILED (lockfileOnlyPackages: 0) ok
vex can't read requirements.txt vex::discover::pypi_other::tests::requirements_files_under_an_ascii_compatible_coding_line_are_read FAILED ok
hosted get wires nothing; vex exit 2; rollback manifest_not_found in_process_rollback_hosted::pypi_requirements_hosted_round_trip_with_an_unmodelled_coding_line FAILED (vex exit 2) ok

Local runs:

  • cargo clippy --workspace --all-features -- -D warnings: clean.
  • rustfmt --check on the touched files: clean. cargo fmt --all -- --check is already red on main across 17 untouched files, and CI doesn't run it.
  • cargo test -p socket-patch-core --all-features --lib: 5889 passed. The other 4 failures are read-only-permission tests that can't fail when run as root in this sandbox (copy_tree relax loop, vlt_heal unremovable lock, poetry/requirements wire-write failure). Untouched code.
  • 13 CLI suites that read requirements files (scan_requirements_lock_only, in_process_rollback_hosted, mode_migration_pypi, in_process_get_hosted_ecosystems, in_process_redirect_pipenv, vendor_eject_fresh_checkout, …): 194 passed. The 1 failure (pipenv_hosted_to_vendored_names_the_unpatched_requirements) needs a live GET to pypi.org, which this sandbox can't reach.
  • The full cargo test --workspace ran out of sandbox disk while linking. CI runs the full matrix.

No wrapper (npm/, pypi/, gem/) changes needed: the decoder is Rust-only.

🤖 Generated with Claude Code

https://claude.ai/code/session_01GwaJm5RXR9MjWKencyZifJ


Note

Low Risk
Narrow decoder fix for template coding headers on ASCII files; behavior for non-ASCII and exotic codecs is unchanged, with broad regression tests.

Overview
Fixes #1212: plain-ASCII requirements.txt files whose PEP 263 header names a codec outside UTF-8 / ASCII / Latin-1 / cp1252 (e.g. iso-8859-15, cp1250) are decoded again instead of being treated as unreadable.

utils::requirements::decode gains Codec::AsciiCompatible, backed by a large alias table (ASCII_COMPATIBLE_CODECS) and python_codec_key normalization so names match Python’s codec lookup. For those codecs, ASCII-only bytes decode like pip; non-ASCII content and codecs that do not map ASCII 1:1 (EBCDIC, UTF-16, unknown names) stay unreadable.

Regression coverage spans the decoder unit test, lock-only scan (root and -r includes), VEX discovery, and hosted get / vex / rollback round-trips.

Reviewed by Cursor Bugbot for commit f1f51da. Configure here.


Generated by Claude Code

Assisted-by: Claude Code:claude-opus-5-5
A plain-ASCII requirements.txt whose PEP 263 header names a codec such
as iso-8859-15, cp1250, mac-roman or gbk installs fine with pip, but
socket-patch reads it as unreadable: scan finds no packages and exits
0, hosted get wires nothing, and vex and rollback fail on an existing
hosted pin.

These tests capture that through the decoder, VEX discovery, lock-only
scan (root file and -r include, hosted and vendored) and the hosted
get / vex / rollback round trip. They fail on main.

Refs #1212

Assisted-by: Claude Code:claude-opus-5-5
A requirements.txt whose PEP 263 header names a codec such as
iso-8859-15, cp1250, mac-roman or gbk was read as unreadable, even
when the file itself is plain ASCII (a header copied from a template).
Scan found no packages and exited 0 while pip installed the unpatched
release; hosted scan and get wired nothing; vex and rollback could not
see an existing hosted pin.

pip decodes with the codec the line names, and every ASCII-compatible
codec decodes ASCII bytes as ASCII. The decoder now knows every Python
name for those codecs (the ISO-8859, Windows, DOS, Mac, KOI8 and CJK
families), looked up the way Python normalizes a codec name, and reads
plain-ASCII files under them. Non-ASCII bytes under those codecs, codecs
that read ASCII differently (EBCDIC, UTF-16/32, UTF-7, HZ, ISO-2022,
escape codecs) and names Python does not know still read as unreadable.

Fixes #1212

Assisted-by: Claude Code:claude-opus-5-5
@mikolalysenko
Mikola Lysenko (mikolalysenko) marked this pull request as ready for review October 9, 2026 03:47
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

BugBot review


Generated by Claude Code

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit f1f51da. Configure here.

@mikolalysenko Mikola Lysenko (mikolalysenko) added the Ready for review Agent-verified: mergeable, CI green, Bugbot clean — awaiting human review label Oct 9, 2026
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

Ready for review at f1f51da.

  • CI: 100/100 check runs green on this head (99 success, 1 skipped).
  • Bugbot: reviewed f1f51da, no findings.
  • Mergeable, no conflicts. Worth a look: the 288-name ASCII-compatible codec table in utils/requirements.rs (generated from Python's codec registry) and the deliberate choice to keep latin-9 unreadable to match pip.

Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

Final review brief

What it does: A requirements file whose PEP 263 header names a codec other than UTF-8, ASCII, Latin-1 or cp1252 (for example # -*- coding: iso-8859-15 -*- copied from a template) was treated as unreadable. Scan reported no packages, hosted get wired nothing, and vex/rollback failed. Now, if the file's bytes are plain ASCII and the named codec decodes ASCII as itself, the file is read as ASCII, which is what pip does. Non-ASCII bytes, codecs that remap ASCII (EBCDIC, UTF-16/32, UTF-7, ISO-2022…) and names Python doesn't know all stay unreadable, as before.

Risk: low. One new arm in coding_line_codec, gated on bytes.is_ascii(), so it can only turn "unreadable" into the exact ASCII text. No other decode path changes.

Look here:

Verified:

  • Read the full diff. It touches 4 files (+216/−2) and CHANGELOG.md is untouched.
  • Checked every codec name in the table against Python 3.11's codecs registry. Every one resolves and decodes all 128 ASCII bytes as ASCII.
  • Confirmed latin-9 and cp-1250 raise LookupError in Python, which matches the PR keeping them unreadable.
  • The new tests cover each reported symptom: the decoder, lock-only scan (root file and a -r include), VEX discovery, and the hosted round trip. The PR shows them red on main at bc0d34c.
  • CI: 477/477 on f1f51da (426 success, 51 skipped), including ci-ok and clippy. Bugbot reviewed f1f51da with no findings. Mergeable, no review threads.

Changes I made: none.

Open questions: none. One non-blocking nit: python_codec_key(&name) is recomputed for each table entry inside the any. This only runs on files with an unmodelled coding line, so it's not a hot path.

Auto-merge (squash) is armed, so approving sends this straight to the merge queue.


Generated by Claude Code

@mikolalysenko
Mikola Lysenko (mikolalysenko) added this pull request to the merge queue Oct 9, 2026
Merged via the queue into main with commit 21536b2 Oct 9, 2026
518 checks passed
@mikolalysenko
Mikola Lysenko (mikolalysenko) deleted the agent/fix-requirements-unmodelled-coding-line branch October 9, 2026 07:19
Mikola Lysenko (mikolalysenko) added a commit that referenced this pull request Oct 9, 2026
…elease notes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Ready for review Agent-verified: mergeable, CI green, Bugbot clean — awaiting human review

Projects

None yet

3 participants