如何使用BeautifulSoup提取XML中指定<p>标签下的文本内容
如何使用BeautifulSoup提取XML中指定
标签下的文本内容
嘿,我来帮你搞定这个问题!你当前的代码是把每个<sec>标签里的所有文本都打印出来,这会把标题和目标段落的内容混在一起,而且没精准定位到你要的<p id="p0055">元素,所以才达不到预期效果。咱们调整下代码,精准定位目标内容就行啦:
方法一:直接定位目标
标签
你可以直接用find()方法,同时指定标签名和id属性,一步找到你要的段落,再提取它的文本:
from bs4 import BeautifulSoup with open('test.xml', 'r') as file: soup = BeautifulSoup(file, 'xml') # 精准匹配id为p0055的<p>标签 target_paragraph = soup.find('p', id='p0055') if target_paragraph: # 提取该标签下的全部文本(包括子标签里的内容,比如上标引用号) print(target_paragraph.get_text())
方法二:先定位父级再找目标
如果你的XML里有多个<sec>标签,担心其他sec里可能有重复id的情况,也可以先定位到目标<sec id="sec2.1">,再在它内部找对应的p标签,这样更稳妥:
from bs4 import BeautifulSoup with open('test.xml', 'r') as file: soup = BeautifulSoup(file, 'xml') # 先找到id为sec2.1的sec标签 target_section = soup.find('sec', id='sec2.1') if target_section: # 在该sec内部找id为p0055的p标签 target_paragraph = target_section.find('p', id='p0055') if target_paragraph: print(target_paragraph.get_text())
为什么你原来的代码没生效?
你之前的代码遍历所有<sec>标签并打印tag.text,这会把每个sec里的所有内容(包括<title>的文本和所有段落文本)全部拼接在一起输出,自然没法单独拿到你要的那段指定p标签里的内容啦。
备注:内容来源于stack exchange,提问作者Alex Maina
相关产品推荐
相关产品推荐

