You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并GDML包含文件并保存,规避解析耗时与字符串长度限制?

Solution for Merging GDML Files Without Full Parsing & Memory Limits

Great question—handling large GDML includes without parsing the entire tree and hitting memory limits is totally doable with a streaming text processing approach. Here's a practical, efficient solution tailored to your needs:

Core Approach

Instead of loading the entire XML tree into memory or relying on full parsing, we'll:

  1. First scan the main GDML file to map <!ENTITY> definitions to their target files (and discard these definitions from the output).
  2. Stream through the main file line-by-line, replacing &entity; references with the actual content of their linked files—writing directly to the output as we go.

This keeps memory usage minimal (only processing one line/file chunk at a time) and avoids the overhead of full XML parsing.

Code Implementation

import re

# Regex patterns to match GDML entity definitions and references
ENTITY_DEF_PATTERN = re.compile(r'^\s*<!ENTITY\s+(\w+)\s+SYSTEM\s+"([^"]+)"\s*>\s*$')
ENTITY_REF_PATTERN = re.compile(r'&(\w+);')

def merge_gdml(main_file_path, output_file_path):
    # Step 1: Collect all entity-to-file mappings from the main GDML
    entity_map = {}
    with open(main_file_path, 'r', encoding='utf-8') as main_file:
        for line in main_file:
            match = ENTITY_DEF_PATTERN.match(line)
            if match:
                entity_name = match.group(1)
                include_file_path = match.group(2)
                entity_map[entity_name] = include_file_path

    # Step 2: Stream process the main file, replacing references and writing output
    with open(main_file_path, 'r', encoding='utf-8') as main_file, \
         open(output_file_path, 'w', encoding='utf-8') as output_file:
        
        for line in main_file:
            # Skip entity definition lines entirely
            if ENTITY_DEF_PATTERN.match(line):
                continue
            
            # Replace any entity references in the current line
            def replace_entity_reference(match):
                entity_name = match.group(1)
                if entity_name not in entity_map:
                    return match.group(0)  # Keep unknown references as-is
                
                # Load and return the content of the included file
                with open(entity_map[entity_name], 'r', encoding='utf-8') as include_file:
                    return include_file.read()
            
            processed_line = ENTITY_REF_PATTERN.sub(replace_entity_reference, line)
            output_file.write(processed_line)

# Example usage
merge_gdml("main.gdml", "merged_output.gdml")

Key Advantages

  • No memory bloat: We never load the entire merged content into a single string—everything is processed line-by-line, so you won't hit sys.maxsize limits.
  • Fast, no full parsing: Regular expressions handle entity matching without parsing the entire XML tree, which is critical for large collections of files.
  • Clean output: Automatically strips out all <!ENTITY> definitions as required.

Handling Nested Includes

If your GDML files have nested include references (i.e., included files have their own <!ENTITY> definitions), you can extend the code to recursively process nested files:

def process_file(file_path, entity_map=None):
    if entity_map is None:
        entity_map = {}
        # First collect entities from this file
        with open(file_path, 'r', encoding='utf-8') as f:
            for line in f:
                match = ENTITY_DEF_PATTERN.match(line)
                if match:
                    entity_map[match.group(1)] = match.group(2)
    
    # Stream process the file, replacing references (including nested ones)
    with open(file_path, 'r', encoding='utf-8') as f:
        for line in f:
            if ENTITY_DEF_PATTERN.match(line):
                continue
            # Recursively replace references with processed include content
            processed_line = ENTITY_REF_PATTERN.sub(
                lambda m: ''.join(process_file(entity_map[m.group(1)], entity_map.copy())),
                line
            )
            yield processed_line

# Updated merge function for nested support
def merge_gdml_with_nested(main_file_path, output_file_path):
    with open(output_file_path, 'w', encoding='utf-8') as output_file:
        for processed_line in process_file(main_file_path):
            output_file.write(processed_line)

Notes

  • Ensure all GDML files use the same encoding (we default to UTF-8; adjust the encoding parameter in open() if needed).
  • The regex patterns assume standard GDML ENTITY formatting—tweak them if your files have non-standard whitespace or syntax.

内容的提问来源于stack exchange,提问作者Keith Sloan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 22:52:36