如何将含章节的XML书籍文件拆分为单个章节XML文件?
拆分XML章节为独立文件的实现方案
方法一:Python 脚本实现
用Python标准库xml.etree.ElementTree即可完成,无需额外依赖:
import xml.etree.ElementTree as ET import os # 解析源XML文件 tree = ET.parse('book.xml') root = tree.getroot() # 创建输出目录(不存在则自动创建) output_dir = 'chapters' os.makedirs(output_dir, exist_ok=True) # 遍历所有chapter节点(若chapter嵌套在其他节点下,改用root.findall('.//chapter')递归查找) for chapter_idx, chapter_node in enumerate(root.findall('chapter'), start=1): # 构建新的XML结构,保留原根节点 new_root = ET.Element(root.tag) new_root.append(chapter_node) # 生成带两位序号的文件名 output_file = os.path.join(output_dir, f'chap{chapter_idx:02d}.xml') # 写入文件,手动添加XML声明保证格式正确 with open(output_file, 'wb') as f: f.write(b'<?xml version="1.0" encoding="UTF-8"?>\n') ET.ElementTree(new_root).write(f, encoding='utf-8') print("章节拆分完成,文件已保存至 chapters 目录")
注意事项
- 如果原XML中
<chapter>节点不是直接在根节点下,把root.findall('chapter')改成root.findall('.//chapter'),递归查找所有层级的章节节点。 - 脚本会自动创建
chapters目录,所有拆分后的文件都会保存在这里。
方法二:XSLT 转换实现
若习惯用XML工具链,可编写XSLT转换规则,用xsltproc或Saxon等工具执行:
1. 编写XSLT转换文件(split_book.xsl)
<?xml version="1.0" encoding="UTF-8"?> <xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:template match="/"> <!-- 遍历所有chapter节点 --> <xsl:for-each select="//chapter"> <!-- 生成带两位序号的输出文件 --> <xsl:result-document method="xml" encoding="UTF-8" href="chap{format-number(position(), '00')}.xml"> <!-- 保留原XML的根节点,并仅保留当前chapter --> <xsl:copy-of select="/*"> <xsl:apply-templates select="/*/*[not(self::chapter)]" mode="remove"/> <xsl:copy-of select="current()"/> </xsl:copy-of> </xsl:result-document> </xsl:for-each> </xsl:template> <!-- 移除根节点下除当前chapter外的其他节点 --> <xsl:template match="*" mode="remove"/> </xsl:stylesheet>
2. 执行转换命令
用xsltproc工具(Linux/macOS默认自带,Windows可安装):
xsltproc split_book.xsl book.xml
执行后会在当前目录生成chap01.xml、chap02.xml等文件。
内容的提问来源于stack exchange,提问作者Umaima Fatima
相关产品推荐
相关产品推荐

