Python提取TXT文件中DHT模块Destination ID对应Huffman表:最后一组1/1表提取失败求助
Hey, I see you're stuck grabbing the final Destination ID = 1 Class = 1 Huffman table block, even though the first three extracted just fine with methods like read() or iterators. The likely culprit is that your current logic uses the ---- separator to detect the end of a block, but the last block doesn't end with another ----—it wraps up right before the *** Marker: SOF0 (Baseline DCT) (xFFC0) *** line. Here are a couple of solid fixes:
1. Use Regular Expressions to Capture All Blocks in One Go
Regex is perfect here because it can handle both end conditions (either a new ---- block start or the final Marker line) without extra logic.
Here's a Python example that grabs every Destination ID block:
import re # Load your entire text file content with open("your_huffman_file.txt", "r") as file: file_content = file.read() # Regex pattern to match each complete Destination ID block block_pattern = r'---- Destination ID = (\d+) Class = (\d+) \(.*?\)(.*?)(?=----|\*\*\* Marker:)' # Use re.DOTALL so . matches newlines, capturing the full block content matches = re.finditer(block_pattern, file_content, re.DOTALL) # Process and output each matched block for match in matches: dest_id = match.group(1) class_id = match.group(2) table_data = match.group(3).strip() print(f"### Destination ID {dest_id} / Class {class_id}") print(table_data) print("---")
How the regex works:
- Starts by matching the block header
---- Destination ID = X Class = Y (...) - Captures all content until it hits either the next
----separator or the*** Marker:line (the end of the last block) re.DOTALLensures the pattern spans multiple lines, so you get the full table content
2. Adjust Line-by-Line Reading to Handle the Final Block
If you prefer line-by-line processing, tweak your termination check to account for the Marker line that ends the last block:
current_block = [] all_blocks = [] with open("your_huffman_file.txt", "r") as file: lines = file.readlines() inside_block = False for line in lines: stripped_line = line.strip() # Check if we're starting a new Destination ID block if stripped_line.startswith("---- Destination ID"): # Save the previous block if we were collecting one if inside_block: all_blocks.append("\n".join(current_block)) current_block = [] inside_block = True current_block.append(stripped_line) # Check if we're ending a block (either new block start or Marker line) elif inside_block: if stripped_line.startswith("----") or stripped_line.startswith("*** Marker:"): all_blocks.append("\n".join(current_block)) current_block = [] inside_block = False # Exit early if we hit the Marker, no more blocks left if stripped_line.startswith("*** Marker:"): break else: current_block.append(stripped_line) # Print out all extracted blocks for i, block in enumerate(all_blocks): print(f"Block {i+1}:") print(block) print("====")
Why Your Original read() Method Failed
Chances are your initial code split the file content using ---- as a delimiter. Since the last block doesn't have a trailing ----, it either got excluded entirely or was merged with the Marker line at the end—this is why the first three worked (they each have a ---- after them) but the last one didn't. Either of the above methods fixes this by accounting for both possible end conditions.
内容的提问来源于stack exchange,提问作者NBG

