Add the four double-byte CJK pages behind an on-by-default dbcs feature
Code pages 932 (Shift-JIS), 936 (GBK), 949 (Unified Hangul), and 950 (Big5), generated from the Unicode Consortium's Microsoft WINDOWS tables. Each page carries a three-state single-byte table (map / lead / undefined) plus sorted pair arrays binary-searched in both directions. Where several codes decode to one scalar (932's NEC/IBM overlap, 950's duplicated box-drawing rows), the encoder's winner is resolved from Microsoft's own bestfit WCTABLE — used only for duplicate resolution, never to import lossy best-fit mappings. The full decode direction was cross-validated against Python's cp932/936/949/950 codecs: zero mismatches over 60k+ pairs. Decode gains TruncatedPair and UndefinedPair failure modes; decode_lossy emits one U+FFFD per failed pair. Both feature configurations are clippy-clean and tested.
This commit is contained in:
25
README.md
25
README.md
@@ -28,22 +28,37 @@ assert_eq!(DosEncoding::Cp874.decode_lossy(&[0xFF]), "\u{FFFD}");
|
||||
|
||||
- **All sixteen single-byte DOS pages** — implemented, decode + encode:
|
||||
437, 737, 775, 850, 852, 855, 857, 860, 861, 862, 863, 864, 865, 866, 869, 874.
|
||||
- **Double-byte CJK pages** (932 Shift-JIS, 936 GBK, 949 Korean, 950 Big5) —
|
||||
planned, to be codegen'd the same way and feature-gated.
|
||||
- **All four double-byte CJK pages** — implemented, decode + encode:
|
||||
932 (Shift-JIS), 936 (GBK), 949 (Unified Hangul), 950 (Big5). They sit
|
||||
behind the on-by-default `dbcs` feature; build with
|
||||
`default-features = false` for a lean single-byte-only crate.
|
||||
|
||||
## Fidelity notes
|
||||
|
||||
- Every table is a verified bijection: anything a page decodes re-encodes to
|
||||
the identical bytes.
|
||||
- Every single-byte table is a verified bijection: anything a page decodes
|
||||
re-encodes to the identical bytes.
|
||||
- Code pages 857, 864, 869, and 874 leave some byte values undefined, so
|
||||
`decode` returns a `Result`; `decode_lossy` substitutes U+FFFD.
|
||||
- Microsoft's CP864 (DOS Arabic) maps `0x25` to ARABIC PERCENT SIGN — the low
|
||||
half is *not* pure ASCII, and `%` itself is unmappable. This is faithful to
|
||||
the source table.
|
||||
- CP932 and CP950 each have a handful of byte pairs that decode to the same
|
||||
character (932's NEC/IBM extension overlap, 950's duplicated box-drawing
|
||||
rows). All of them decode; re-encoding picks the code Windows picks, as
|
||||
resolved from Microsoft's own `bestfit932/950.txt` WCTABLE (vendored in
|
||||
`data/`, used *only* for duplicate resolution — no lossy best-fit mappings
|
||||
are imported).
|
||||
- Double-byte decode adds two failure modes: a lead byte at end of input
|
||||
(`TruncatedPair`) and an undefined lead/trail pair (`UndefinedPair`).
|
||||
`decode_lossy` emits one U+FFFD per failed pair.
|
||||
- The full decode direction of all four CJK pages agrees exactly with
|
||||
Python's `cp932`/`cp936`/`cp949`/`cp950` codecs (59k+ pairs, verified
|
||||
during development).
|
||||
|
||||
## Regenerating tables
|
||||
|
||||
`src/tables.rs` is generated from the vendored mapping files in `data/`
|
||||
Everything in `src/tables/` except `mod.rs` is generated from the vendored
|
||||
mapping files in `data/`
|
||||
(needs Python ≥ 3.10 and `rustfmt` on `PATH`):
|
||||
|
||||
```sh
|
||||
|
||||
Reference in New Issue
Block a user