You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析DITA Map及递归提取引用文件 求助可行脚本

递归解析DITA Map/XML文件的Python脚本

以下是一个可直接使用的脚本,能递归遍历DITA Map/XML文件及其引用的所有关联文件,最终生成完整的文件列表:

import os
import xml.etree.ElementTree as ET

# 定义需要收集的文件扩展名
TARGET_EXTENSIONS = ('.dita', '.ditamap', '.xml')
# 记录已处理的文件,避免循环引用和重复处理
processed_files = set()

def parse_dita_file(file_path):
    # 标准化文件路径,消除路径格式差异
    normalized_path = os.path.normpath(os.path.abspath(file_path))
    
    # 跳过已处理的文件
    if normalized_path in processed_files:
        return []
    processed_files.add(normalized_path)
    
    print(f"正在处理: {normalized_path}")
    file_list = [normalized_path]
    
    try:
        # 解析XML文件
        tree = ET.parse(normalized_path)
        root = tree.getroot()
        
        # 自动提取DITA文件的命名空间
        namespace = ''
        if root.tag.startswith('{'):
            namespace = root.tag.split('}')[0] + '}'
        
        # 查找所有带href属性的引用元素(覆盖常见的topicref、mapref)
        ref_elements = root.findall(f'.//{namespace}topicref') + root.findall(f'.//{namespace}mapref')
        
        for elem in ref_elements:
            href = elem.get('href')
            if not href:
                continue
            
            # 去除href中的锚点部分(如file.dita#topic-id)
            if '#' in href:
                href = href.split('#')[0]
            
            # 过滤不符合扩展名要求的文件
            if not href.lower().endswith(TARGET_EXTENSIONS):
                continue
            
            # 拼接引用文件的绝对路径
            ref_file_path = os.path.join(os.path.dirname(normalized_path), href)
            # 递归处理引用文件并合并结果
            file_list.extend(parse_dita_file(ref_file_path))
    
    except Exception as e:
        print(f"处理文件 {normalized_path} 出错: {str(e)}")
    
    return file_list

if __name__ == '__main__':
    # 替换为你的起始DITA Map/XML文件路径
    start_file = 'path/to/your/start.ditamap'
    if not os.path.exists(start_file):
        print(f"错误:文件 {start_file} 不存在")
        exit(1)
    
    # 执行解析
    all_files = parse_dita_file(start_file)
    
    # 输出最终文件列表
    print("\n===== 所有关联文件列表 =====")
    for file in all_files:
        print(file)
    
    # 将列表写入文件,方便后续打包使用
    with open('dita_files_list.txt', 'w', encoding='utf-8') as f:
        f.write('\n'.join(all_files))
    print("\n文件列表已保存到 dita_files_list.txt")

关键功能说明

  • 路径标准化:用os.path.abspath和os.path.normpath统一路径格式,避免因相对路径、大小写问题重复记录文件。
  • 命名空间兼容:自动识别DITA文件的XML命名空间,确保能正确定位<topicref>、<mapref>等引用元素。
  • 循环引用防护:通过processed_files集合记录已处理文件,防止无限递归和重复条目。
  • href处理:自动剥离锚点部分,只保留实际文件名。
  • 容错机制:捕获解析异常,单个文件出错不会导致整个脚本终止。

使用方法

  1. 将start_file变量替换为你的起始DITA Map/XML文件路径。
  2. 运行脚本,会实时打印正在处理的文件,最终输出完整列表并保存到dita_files_list.txt。
  3. 若你的DITA文件包含其他引用元素(如<xref>),可在ref_elements的查找语句中添加对应的标签。

内容的提问来源于stack exchange,提问作者Russ Urquhart

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 23:55:20