BeautifulSoup如何排除p标签下嵌套的指定img元素?
问题说明
做网页数据抓取时,需要提取目标页面所有h1、p标签内容输出为HTML文件,页面仅包含1个h1标签,存在多个p标签。初始代码提取时,会把p标签内嵌套的返回顶部按钮img元素(DOM路径为p.a.img)一并提取,对最终输出造成干扰。
初始实现代码如下:
from bs4 import BeautifulSoup as bs import requests url = 'https://chhouk-krohom.com/%E1%9E%91%E1%9E%B8%E1%9E%83%E1%9E%93%E1%9E%B7%E1%9E%80%E1%9E%B6%E1%9E%99%E1%9F%A1%E1%9F%A4/' response = requests.get(url) soup = bs(response.content, 'html.parser') contents = soup.find_all(['h1', 'p']) for content in contents: print(content) content = soup.prettify() with open('sutta.html', 'wt', encoding='utf-8') as file: file.write(str(content))
解决方法
遍历提取到的h1、p标签时,递归查找标签下所有img元素,直接从DOM结构中移除即可,修改后代码如下:
from bs4 import BeautifulSoup as bs import requests url = 'https://chhouk-krohom.com/%E1%9E%91%E1%9E%B8%E1%9E%83%E1%9E%93%E1%9E%B7%E1%9E%80%E1%9E%B6%E1%9E%99%E1%9F%A1%E1%9F%A4/' response = requests.get(url) soup = bs(response.content, 'html.parser') contents = soup.find_all(['h1', 'p']) output_soup = bs('', 'html.parser') for content in contents: # 递归删除当前标签下所有层级的img元素 for img in content.find_all('img', recursive=True): img.decompose() output_soup.append(content) with open('sutta.html', 'wt', encoding='utf-8') as file: file.write(output_soup.prettify())
关键逻辑说明
content.find_all('img', recursive=True)会匹配当前标签下任意嵌套层级的img元素,不管是p标签直接子元素,还是藏在a标签等子节点内的img都能被找到decompose()方法会将匹配到的元素从DOM树中彻底移除,不会残留多余标签- 单独构建输出用的soup对象存储过滤后的内容,避免把原页面中不需要的其他节点写入最终HTML文件
内容的提问来源于stack exchange,提问作者Mortus Pect
相关产品推荐
相关产品推荐

