You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取TXT文件中DHT模块Destination ID对应Huffman表:最后一组1/1表提取失败求助

Trouble Extracting the Last Destination ID Huffman Table Block

Hey, I see you're stuck grabbing the final Destination ID = 1 Class = 1 Huffman table block, even though the first three extracted just fine with methods like read() or iterators. The likely culprit is that your current logic uses the ---- separator to detect the end of a block, but the last block doesn't end with another ----—it wraps up right before the *** Marker: SOF0 (Baseline DCT) (xFFC0) *** line. Here are a couple of solid fixes:

1. Use Regular Expressions to Capture All Blocks in One Go

Regex is perfect here because it can handle both end conditions (either a new ---- block start or the final Marker line) without extra logic.

Here's a Python example that grabs every Destination ID block:

import re

# Load your entire text file content
with open("your_huffman_file.txt", "r") as file:
    file_content = file.read()

# Regex pattern to match each complete Destination ID block
block_pattern = r'---- Destination ID = (\d+) Class = (\d+) \(.*?\)(.*?)(?=----|\*\*\* Marker:)'
# Use re.DOTALL so . matches newlines, capturing the full block content
matches = re.finditer(block_pattern, file_content, re.DOTALL)

# Process and output each matched block
for match in matches:
    dest_id = match.group(1)
    class_id = match.group(2)
    table_data = match.group(3).strip()
    
    print(f"### Destination ID {dest_id} / Class {class_id}")
    print(table_data)
    print("---")

How the regex works:

  • Starts by matching the block header ---- Destination ID = X Class = Y (...)
  • Captures all content until it hits either the next ---- separator or the *** Marker: line (the end of the last block)
  • re.DOTALL ensures the pattern spans multiple lines, so you get the full table content

2. Adjust Line-by-Line Reading to Handle the Final Block

If you prefer line-by-line processing, tweak your termination check to account for the Marker line that ends the last block:

current_block = []
all_blocks = []

with open("your_huffman_file.txt", "r") as file:
    lines = file.readlines()
    inside_block = False
    
    for line in lines:
        stripped_line = line.strip()
        
        # Check if we're starting a new Destination ID block
        if stripped_line.startswith("---- Destination ID"):
            # Save the previous block if we were collecting one
            if inside_block:
                all_blocks.append("\n".join(current_block))
                current_block = []
            inside_block = True
            current_block.append(stripped_line)
        
        # Check if we're ending a block (either new block start or Marker line)
        elif inside_block:
            if stripped_line.startswith("----") or stripped_line.startswith("*** Marker:"):
                all_blocks.append("\n".join(current_block))
                current_block = []
                inside_block = False
                # Exit early if we hit the Marker, no more blocks left
                if stripped_line.startswith("*** Marker:"):
                    break
            else:
                current_block.append(stripped_line)

# Print out all extracted blocks
for i, block in enumerate(all_blocks):
    print(f"Block {i+1}:")
    print(block)
    print("====")

Why Your Original read() Method Failed

Chances are your initial code split the file content using ---- as a delimiter. Since the last block doesn't have a trailing ----, it either got excluded entirely or was merged with the Marker line at the end—this is why the first three worked (they each have a ---- after them) but the last one didn't. Either of the above methods fixes this by accounting for both possible end conditions.

内容的提问来源于stack exchange,提问作者NBG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 00:33:10