如何合并GDML包含文件并保存,规避解析耗时与字符串长度限制?
Solution for Merging GDML Files Without Full Parsing & Memory Limits
Great question—handling large GDML includes without parsing the entire tree and hitting memory limits is totally doable with a streaming text processing approach. Here's a practical, efficient solution tailored to your needs:
Core Approach
Instead of loading the entire XML tree into memory or relying on full parsing, we'll:
- First scan the main GDML file to map
<!ENTITY>definitions to their target files (and discard these definitions from the output). - Stream through the main file line-by-line, replacing
&entity;references with the actual content of their linked files—writing directly to the output as we go.
This keeps memory usage minimal (only processing one line/file chunk at a time) and avoids the overhead of full XML parsing.
Code Implementation
import re # Regex patterns to match GDML entity definitions and references ENTITY_DEF_PATTERN = re.compile(r'^\s*<!ENTITY\s+(\w+)\s+SYSTEM\s+"([^"]+)"\s*>\s*$') ENTITY_REF_PATTERN = re.compile(r'&(\w+);') def merge_gdml(main_file_path, output_file_path): # Step 1: Collect all entity-to-file mappings from the main GDML entity_map = {} with open(main_file_path, 'r', encoding='utf-8') as main_file: for line in main_file: match = ENTITY_DEF_PATTERN.match(line) if match: entity_name = match.group(1) include_file_path = match.group(2) entity_map[entity_name] = include_file_path # Step 2: Stream process the main file, replacing references and writing output with open(main_file_path, 'r', encoding='utf-8') as main_file, \ open(output_file_path, 'w', encoding='utf-8') as output_file: for line in main_file: # Skip entity definition lines entirely if ENTITY_DEF_PATTERN.match(line): continue # Replace any entity references in the current line def replace_entity_reference(match): entity_name = match.group(1) if entity_name not in entity_map: return match.group(0) # Keep unknown references as-is # Load and return the content of the included file with open(entity_map[entity_name], 'r', encoding='utf-8') as include_file: return include_file.read() processed_line = ENTITY_REF_PATTERN.sub(replace_entity_reference, line) output_file.write(processed_line) # Example usage merge_gdml("main.gdml", "merged_output.gdml")
Key Advantages
- No memory bloat: We never load the entire merged content into a single string—everything is processed line-by-line, so you won't hit
sys.maxsizelimits. - Fast, no full parsing: Regular expressions handle entity matching without parsing the entire XML tree, which is critical for large collections of files.
- Clean output: Automatically strips out all
<!ENTITY>definitions as required.
Handling Nested Includes
If your GDML files have nested include references (i.e., included files have their own <!ENTITY> definitions), you can extend the code to recursively process nested files:
def process_file(file_path, entity_map=None): if entity_map is None: entity_map = {} # First collect entities from this file with open(file_path, 'r', encoding='utf-8') as f: for line in f: match = ENTITY_DEF_PATTERN.match(line) if match: entity_map[match.group(1)] = match.group(2) # Stream process the file, replacing references (including nested ones) with open(file_path, 'r', encoding='utf-8') as f: for line in f: if ENTITY_DEF_PATTERN.match(line): continue # Recursively replace references with processed include content processed_line = ENTITY_REF_PATTERN.sub( lambda m: ''.join(process_file(entity_map[m.group(1)], entity_map.copy())), line ) yield processed_line # Updated merge function for nested support def merge_gdml_with_nested(main_file_path, output_file_path): with open(output_file_path, 'w', encoding='utf-8') as output_file: for processed_line in process_file(main_file_path): output_file.write(processed_line)
Notes
- Ensure all GDML files use the same encoding (we default to UTF-8; adjust the
encodingparameter inopen()if needed). - The regex patterns assume standard GDML
ENTITYformatting—tweak them if your files have non-standard whitespace or syntax.
内容的提问来源于stack exchange,提问作者Keith Sloan
相关产品推荐
相关产品推荐

