You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用lxml解析自定义XML并生成符合规则的字符串列表?

高效转换自定义XML为指定字符串列表的lxml实现

问题概述

给定自定义XML结构,需转换为满足以下规则的字符串列表:

  • 移除<b>标签,保留其文本内容
  • 每个<heading>和<p>初始作为独立元素,末尾无句号则自动补充
  • 若段落文本以冒号引出<ul>列表,将列表项以- 前缀换行整合到当前段落
  • 若段落文本无冒号直接接<ul>列表,将段落拆分为「前缀文本」「列表项」「后缀文本」三个独立元素

输入XML

<?xml version="1.0"?>
<body>
    <heading><b>This is a title</b></heading>
    <p>This is a first <b>paragraph</b>.</p>
    <p>This is a second <b>paragraph</b>. With a list: 
        <ul>
            <li>first item</li>
            <li>second item</li>
        </ul>
    And the end.
    </p>
    <p>This is a third paragraph.
        <ul>
            <li>This is a first long sentence.</li>
            <li>This is a second long sentence.</li>
        </ul>
    And the end of the paragraph.</p>
</body>

期望输出

[
    "This is a title.", # Note the period
    "This is a first paragraph.",
    "This is a second paragraph. With a list:\n- first item\n- second item\nAnd the end.",
    "This is a third paragraph.",
    "This is a first long sentence.",
    "This is a second long sentence.",
    "And the end of the paragraph."
]

实现方案

利用lxml的节点遍历和文本提取能力,通过XPath精准定位节点,针对不同场景处理列表整合与拆分:

from lxml import etree

def parse_xml_to_list(xml_str):
    root = etree.fromstring(xml_str)
    output = []
    
    # 遍历所有heading和p节点
    for node in root.xpath('heading | p'):
        # 提取节点内所有文本(自动忽略b标签),清理多余空白
        raw_text_parts = [text.strip() for text in node.itertext() if text.strip()]
        # 分离前缀文本(排除列表项文本)
        ul_node = node.find('ul')
        if ul_node:
            list_item_texts = [li.text.strip() for li in ul_node.findall('li')]
            prefix_text = ' '.join([part for part in raw_text_parts if part not in list_item_texts])
        else:
            prefix_text = ' '.join(raw_text_parts)
        
        if not ul_node:
            # 无列表节点:处理句号后添加到结果
            final_text = prefix_text if prefix_text.endswith('.') else f"{prefix_text}."
            output.append(final_text)
            continue
        
        # 提取列表后的文本
        suffix_text = ' '.join([text.strip() for text in node.xpath('ul/following-sibling::text()') if text.strip()])
        list_items = [li.text.strip() for li in ul_node.findall('li')]
        
        if prefix_text.endswith(':'):
            # 冒号引出列表:整合为单一元素
            formatted_list = '\n- ' + '\n- '.join(list_items)
            full_text = f"{prefix_text}{formatted_list}"
            if suffix_text:
                full_text += f"\n{suffix_text}"
            # 补充句号(如果末尾没有)
            if not full_text.endswith('.'):
                full_text += '.'
            output.append(full_text)
        else:
            # 无冒号:拆分元素
            # 添加前缀文本
            if prefix_text:
                output.append(prefix_text if prefix_text.endswith('.') else f"{prefix_text}.")
            # 添加列表项
            for item in list_items:
                output.append(item if item.endswith('.') else f"{item}.")
            # 添加后缀文本
            if suffix_text:
                output.append(suffix_text if suffix_text.endswith('.') else f"{suffix_text}.")
    
    return output

# 测试执行
xml_input = '''<?xml version="1.0"?>
<body>
    <heading><b>This is a title</b></heading>
    <p>This is a first <b>paragraph</b>.</p>
    <p>This is a second <b>paragraph</b>. With a list: 
        <ul>
            <li>first item</li>
            <li>second item</li>
        </ul>
    And the end.
    </p>
    <p>This is a third paragraph.
        <ul>
            <li>This is a first long sentence.</li>
            <li>This is a second long sentence.</li>
        </ul>
    And the end of the paragraph.</p>
</body>'''

result = parse_xml_to_list(xml_input)
print(result)

关键逻辑说明

  1. 文本提取:itertext()自动遍历节点所有子文本,无需手动处理<b>标签,配合列表推导式清理空白。
  2. 列表检测:通过find('ul')快速判断当前段落是否包含列表,避免冗余遍历。
  3. 场景分支:根据前缀文本是否以冒号结尾,分别执行「整合列表」或「拆分段落」逻辑。
  4. 句号处理:统一对所有文本片段检查末尾标点,确保符合规则。

内容的提问来源于stack exchange,提问作者Vincent

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 16:50:29