Bidirectional, emulator-grade conversion between classic Mac OS text encodings and Unicode. Renamed at extraction-to-shared-crate time (2026-07-19): consumers will be fsinspect, adx, and ad-decoder via git dependencies on code.movq.us.
43 lines
1.4 KiB
Markdown
43 lines
1.4 KiB
Markdown
# apple-encodings
|
|
|
|
Bidirectional, emulator-grade conversion between classic Mac OS text encodings
|
|
and Unicode — without linking ICU. The crate owns the canonical Apple tables, so
|
|
it stays small, self-contained, and cross-compiles cleanly.
|
|
|
|
```rust
|
|
use apple_encodings::{AppleEncoding, MacRomanRevision};
|
|
|
|
let enc = AppleEncoding::default(); // Mac OS Roman, post-8.5
|
|
assert_eq!(enc.decode(b"Caf\x8e"), "Café");
|
|
assert_eq!(enc.encode("Café").unwrap(), b"Caf\x8e");
|
|
|
|
// Pick by the Finder Info `fdScript` byte:
|
|
let enc = AppleEncoding::from_script_code(apple_encodings::SCRIPT_ROMAN).unwrap();
|
|
|
|
// Revision matters for exactly one byte (0xDB):
|
|
let classic = AppleEncoding::MacRoman(MacRomanRevision::Classic);
|
|
assert_eq!(classic.decode(&[0xDB]), "¤"); // pre-8.5 currency sign, not €
|
|
```
|
|
|
|
## Status
|
|
|
|
- **Mac OS Roman** — implemented, both pre- and post-8.5 revisions, decode + encode.
|
|
- **Regional single-byte** (Cyrillic, Greek, Turkish, …) and **CJK double-byte**
|
|
(Japanese, Big5, GB, Korean) — planned, to be codegen'd from the Unicode
|
|
Consortium `VENDORS/APPLE/*.TXT` tables. The double-byte tables will be
|
|
feature-gated.
|
|
|
|
## Scope
|
|
|
|
Text-encoding conversion only. Unicode normalization concerns (e.g. HFS+'s
|
|
NFD-ish decomposition) are deliberately **out of scope** and left to consumers.
|
|
|
|
## Consumers
|
|
|
|
Designed to be shared by `ad-decoder`/`adx` and `fsinspect` via path (and later
|
|
git) dependency.
|
|
|
|
## License
|
|
|
|
MIT
|