WOFF2目录项标志(DirectoryEntry Flags)解析异常问题排查
Hey there! Let's dig into why you're seeing duplicate table names in your WOFF2 parser—this is a tricky spot, but let's break down the likely issues and fixes.
First, let's confirm: your flag handling for the table index is actually correct. Both (flags << 2) >> 2 and flags & 0x3f will isolate the first 6 bits (since 0x3f is 00111111 in binary), so that part isn't the problem. The duplicate names are coming from somewhere else.
Here are the top troubleshooting directions to check:
1. Verify Your Table Name List is Built Correctly
The table name list comes right after the WOFF2 header, and it's a sequence of 4-byte ASCII strings (one per table). If this list is parsed wrong, your index will point to the wrong name (or duplicates).
- Double-check that you're reading exactly
numTables * 4bytes for this list (wherenumTablescomes from the WOFF2 header's 16-bit big-endian value). - Print out your
tablesarray and compare it against a known-good WOFF2 file (use a tool likewoff2_decompressto extract the OTF and check its table list). If the list has duplicate or garbled names here, that's your root cause.
2. Fix Byte Order for Multi-Byte Fields
While Flags is a single byte (no endianness issues), almost every other field in WOFF2 uses big-endian (network byte order). If your parsing tool defaults to little-endian, you'll read incorrect values that throw off your entire parsing flow:
- Header fields like
numTables,flavor, andlengthmust be read as big-endian 16/32-bit integers. - Table dictionary entries have 3-byte big-endian
LengthandOffsetvalues. If you're reading these as 4-byte little-endian integers, you'll jump to the wrong position in the file, leading to misread flags and duplicate names.
3. Ensure You're Reading Table Dictionary Entries Properly
Each table dictionary entry is exactly 8 bytes long, per the WOFF2 spec:
- 1 byte Flags
- 3 bytes Length (big-endian)
- 3 bytes Offset (big-endian)
- 1 byte Reserved (must be 0)
- If you're skipping the reserved byte, or reading Length/Offset as 4 bytes instead of 3, you'll shift all subsequent entries out of alignment. This means you'll be reading garbage bytes as flags, leading to wrong table indices and duplicates.
- Add a check for the reserved bit in Flags (bit 7,
flags & 0x80). If this bit is ever set, you know your entry alignment is off—since the spec requires this bit to be 0.
4. Validate Your Entry Count Matches numTables
Make sure you're reading exactly numTables entries from the table dictionary. If you loop one too many times, you'll start reading data from the next section (like table data) as flags, which will give you random indices and duplicate names.
Here's a quick adjusted code snippet to implement these checks:
// Read numTables from header (big-endian 16-bit) let numTables = binary.getUInt16BE() // Build valid table name list var tables: [String] = [] for _ in 0..<numTables { let nameBytes = binary.getBytes(4) guard let tableName = String(bytes: nameBytes, encoding: .ascii) else { fatalError("Invalid table name bytes") } tables.append(tableName) } // Parse each table dictionary entry correctly for _ in 0..<numTables { let flags = binary.getUInt8() // Check reserved bit is 0 (spec requirement) guard (flags & 0x80) == 0 else { fatalError("Corrupted flags: reserved bit is set") } let tableNameIndex = flags & 0x3f guard tableNameIndex < tables.count else { fatalError("Table index \(tableNameIndex) out of bounds") } let tableName = tables[Int(tableNameIndex)] // Read 3-byte big-endian Length let lengthBytes = binary.getBytes(3) let length = (UInt32(lengthBytes[0]) << 16) | (UInt32(lengthBytes[1]) << 8) | UInt32(lengthBytes[2]) // Read 3-byte big-endian Offset let offsetBytes = binary.getBytes(3) let offset = (UInt32(offsetBytes[0]) << 16) | (UInt32(offsetBytes[1]) << 8) | UInt32(offsetBytes[2]) // Skip reserved byte _ = binary.getUInt8() // Process table data... }
Start with verifying the table name list and entry alignment—those are the most common culprits here.
内容的提问来源于stack exchange,提问作者gustav

