使用Python ElementTree提取XML中content内的img src属性值
问题:从Atom XML的content节点中提取img的src属性值
我能轻松获取Atom XML里的title和link,但不知道怎么从content节点中提取img的src属性值。
XML示例
<feed> <entry> <title>A title</title> <link rel="alternate" type="text/html" href="https://url.html"/> <content> <figure> <img alt="text" src="https://image.jpg" /> <p>some text</p></content> </entry> </feed>
当前使用的Python代码
for entry in root.findall('{http://www.w3.org/2005/Atom}entry'): title = entry.find('{http://www.w3.org/2005/Atom}title').text link = entry.find('{http://www.w3.org/2005/Atom}link').attrib['href'] img_url = entry.find('{http://www.w3.org/2005/Atom}src') arr.append([title, link, img_url]) print(title, link) return render(request, 'madxmlparser.html', {'arr':arr})
解决方案
你当前的代码直接查找src节点是错误的——src是img标签的属性,且content节点内的内容是HTML格式的文本,并非XML子节点。需要先提取content的文本内容,再用HTML解析工具提取img的src属性。
步骤:
- 安装并导入
BeautifulSoup库(未安装的话执行命令:pip install beautifulsoup4) - 获取
content节点的文本内容 - 用
BeautifulSoup解析这段HTML,定位img标签并提取src属性
修改后的代码
from bs4 import BeautifulSoup # 假设root是已解析的XML根节点 for entry in root.findall('{http://www.w3.org/2005/Atom}entry'): title = entry.find('{http://www.w3.org/2005/Atom}title').text link = entry.find('{http://www.w3.org/2005/Atom}link').attrib['href'] # 获取content节点的文本内容,做空值容错 content_elem = entry.find('{http://www.w3.org/2005/Atom}content') content_text = content_elem.text if content_elem else '' # 解析HTML内容 soup = BeautifulSoup(content_text, 'html.parser') # 提取第一个img标签的src,做空值容错 img_tag = soup.find('img') img_url = img_tag['src'] if img_tag else None arr.append([title, link, img_url]) print(title, link, img_url) return render(request, 'madxmlparser.html', {'arr':arr})
补充说明:
- 如果
content内有多个img标签,可使用soup.find_all('img')获取所有标签,再遍历提取每个的src属性 - 代码中加入了空值判断,避免因content节点不存在或无img标签导致程序报错
内容的提问来源于stack exchange,提问作者metalboxhead
相关产品推荐
相关产品推荐

