如何用lxml解析自定义XML并生成符合规则的字符串列表?
高效转换自定义XML为指定字符串列表的lxml实现
问题概述
给定自定义XML结构,需转换为满足以下规则的字符串列表:
- 移除
<b>标签,保留其文本内容 - 每个
<heading>和<p>初始作为独立元素,末尾无句号则自动补充 - 若段落文本以冒号引出
<ul>列表,将列表项以-前缀换行整合到当前段落 - 若段落文本无冒号直接接
<ul>列表,将段落拆分为「前缀文本」「列表项」「后缀文本」三个独立元素
输入XML
<?xml version="1.0"?> <body> <heading><b>This is a title</b></heading> <p>This is a first <b>paragraph</b>.</p> <p>This is a second <b>paragraph</b>. With a list: <ul> <li>first item</li> <li>second item</li> </ul> And the end. </p> <p>This is a third paragraph. <ul> <li>This is a first long sentence.</li> <li>This is a second long sentence.</li> </ul> And the end of the paragraph.</p> </body>
期望输出
[ "This is a title.", # Note the period "This is a first paragraph.", "This is a second paragraph. With a list:\n- first item\n- second item\nAnd the end.", "This is a third paragraph.", "This is a first long sentence.", "This is a second long sentence.", "And the end of the paragraph." ]
实现方案
利用lxml的节点遍历和文本提取能力,通过XPath精准定位节点,针对不同场景处理列表整合与拆分:
from lxml import etree def parse_xml_to_list(xml_str): root = etree.fromstring(xml_str) output = [] # 遍历所有heading和p节点 for node in root.xpath('heading | p'): # 提取节点内所有文本(自动忽略b标签),清理多余空白 raw_text_parts = [text.strip() for text in node.itertext() if text.strip()] # 分离前缀文本(排除列表项文本) ul_node = node.find('ul') if ul_node: list_item_texts = [li.text.strip() for li in ul_node.findall('li')] prefix_text = ' '.join([part for part in raw_text_parts if part not in list_item_texts]) else: prefix_text = ' '.join(raw_text_parts) if not ul_node: # 无列表节点:处理句号后添加到结果 final_text = prefix_text if prefix_text.endswith('.') else f"{prefix_text}." output.append(final_text) continue # 提取列表后的文本 suffix_text = ' '.join([text.strip() for text in node.xpath('ul/following-sibling::text()') if text.strip()]) list_items = [li.text.strip() for li in ul_node.findall('li')] if prefix_text.endswith(':'): # 冒号引出列表:整合为单一元素 formatted_list = '\n- ' + '\n- '.join(list_items) full_text = f"{prefix_text}{formatted_list}" if suffix_text: full_text += f"\n{suffix_text}" # 补充句号(如果末尾没有) if not full_text.endswith('.'): full_text += '.' output.append(full_text) else: # 无冒号:拆分元素 # 添加前缀文本 if prefix_text: output.append(prefix_text if prefix_text.endswith('.') else f"{prefix_text}.") # 添加列表项 for item in list_items: output.append(item if item.endswith('.') else f"{item}.") # 添加后缀文本 if suffix_text: output.append(suffix_text if suffix_text.endswith('.') else f"{suffix_text}.") return output # 测试执行 xml_input = '''<?xml version="1.0"?> <body> <heading><b>This is a title</b></heading> <p>This is a first <b>paragraph</b>.</p> <p>This is a second <b>paragraph</b>. With a list: <ul> <li>first item</li> <li>second item</li> </ul> And the end. </p> <p>This is a third paragraph. <ul> <li>This is a first long sentence.</li> <li>This is a second long sentence.</li> </ul> And the end of the paragraph.</p> </body>''' result = parse_xml_to_list(xml_input) print(result)
关键逻辑说明
- 文本提取:
itertext()自动遍历节点所有子文本,无需手动处理<b>标签,配合列表推导式清理空白。 - 列表检测:通过
find('ul')快速判断当前段落是否包含列表,避免冗余遍历。 - 场景分支:根据前缀文本是否以冒号结尾,分别执行「整合列表」或「拆分段落」逻辑。
- 句号处理:统一对所有文本片段检查末尾标点,确保符合规则。
内容的提问来源于stack exchange,提问作者Vincent
相关产品推荐
相关产品推荐

