You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取XML中指定<p>标签下的文本内容

如何使用BeautifulSoup提取XML中指定

标签下的文本内容

嘿,我来帮你搞定这个问题!你当前的代码是把每个<sec>标签里的所有文本都打印出来,这会把标题和目标段落的内容混在一起,而且没精准定位到你要的<p id="p0055">元素,所以才达不到预期效果。咱们调整下代码,精准定位目标内容就行啦:

方法一:直接定位目标

标签

你可以直接用find()方法,同时指定标签名和id属性,一步找到你要的段落,再提取它的文本:

from bs4 import BeautifulSoup

with open('test.xml', 'r') as file:
    soup = BeautifulSoup(file, 'xml')

# 精准匹配id为p0055的<p>标签
target_paragraph = soup.find('p', id='p0055')

if target_paragraph:
    # 提取该标签下的全部文本(包括子标签里的内容,比如上标引用号)
    print(target_paragraph.get_text())

方法二:先定位父级再找目标

如果你的XML里有多个<sec>标签,担心其他sec里可能有重复id的情况,也可以先定位到目标<sec id="sec2.1">,再在它内部找对应的p标签,这样更稳妥:

from bs4 import BeautifulSoup

with open('test.xml', 'r') as file:
    soup = BeautifulSoup(file, 'xml')

# 先找到id为sec2.1的sec标签
target_section = soup.find('sec', id='sec2.1')
if target_section:
    # 在该sec内部找id为p0055的p标签
    target_paragraph = target_section.find('p', id='p0055')
    if target_paragraph:
        print(target_paragraph.get_text())

为什么你原来的代码没生效?

你之前的代码遍历所有<sec>标签并打印tag.text,这会把每个sec里的所有内容(包括<title>的文本和所有段落文本)全部拼接在一起输出,自然没法单独拿到你要的那段指定p标签里的内容啦。

备注:内容来源于stack exchange,提问作者Alex Maina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 13:59:32