You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取XML标签文本并生成指定格式字典?

使用BeautifulSoup实现XML转目标字典的最优方法

最优实现方式是利用字典推导式结合BeautifulSoup的文本提取方法,代码简洁高效,逻辑清晰:

完整代码示例

from bs4 import BeautifulSoup

# 待处理的XML内容
xml_content = """
<article id = '1'> 
  <p> This is </p> 
  <p> example A </p>
</article>

<article id = '2'> 
  <p> This is </p> 
  <p> example B </p>
</article>
"""

# 解析XML,推荐使用lxml解析器(需提前安装:pip install lxml)
soup = BeautifulSoup(xml_content, "lxml")

# 生成目标字典
result_dict = {
    int(article['id']): ' '.join(article.stripped_strings)
    for article in soup.find_all('article')
}

print(result_dict)
# 输出:{1: 'This is example A', 2: 'This is example B'}

关键逻辑说明

  • soup.find_all('article'):批量获取所有<article>标签节点
  • int(article['id']):提取标签的id属性并转为整数,作为字典的键
  • ' '.join(article.stripped_strings):自动提取当前<article>下所有文本内容,同时剔除每个文本段的前后空白、换行符,再用空格拼接成完整字符串,完美解决多<p>标签的文本合并需求

这种写法既避免了冗余的循环代码,又充分利用了BeautifulSoup的内置方法处理文本格式,是实现该需求的最优方案。

内容的提问来源于stack exchange,提问作者Dieu94

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 18:30:59