使用Beautiful Soup提取网页文本:去重复链接并分离脚注
解决网页文本提取的重复链接与脚注分离问题
针对你的需求,我们可以通过定位特定页面容器和过滤无关标签来解决重复链接问题,同时单独提取脚注区域实现内容分离。以下是修改后的代码及说明:
修改后的完整代码
from bs4 import BeautifulSoup from bs4.element import Comment import urllib.request def tag_visible(element): # 排除样式、脚本、头部、导航、页脚等无关标签的内容 if element.parent.name in ['style', 'script', 'head', 'title', 'meta', '[document]', 'nav', 'footer', 'header']: return False # 排除注释内容 if isinstance(element, Comment): return False return True def extract_main_content_and_footnotes(body): soup = BeautifulSoup(body, 'html.parser') # 1. 提取正文:定位到页面正文所在的容器(针对目标网页的结构) main_content_container = soup.find('div', class_='entry-content') # 如果找不到指定容器, fallback 到整个body if not main_content_container: main_content_container = soup.body # 过滤正文里的可见文本,去除空内容 main_texts = main_content_container.find_all(text=True) visible_main_texts = filter(tag_visible, main_texts) main_text = " ".join(t.strip() for t in visible_main_texts if t.strip()) # 2. 提取脚注:定位到页面脚注所在的容器 footnotes_container = soup.find('div', class_='footnotes') footnotes_text = "" if footnotes_container: footnotes_texts = footnotes_container.find_all(text=True) visible_footnotes = filter(tag_visible, footnotes_texts) footnotes_text = " ".join(t.strip() for t in visible_footnotes if t.strip()) return main_text, footnotes_text # 调用函数提取内容 html = urllib.request.urlopen('https://ordoabchao.ca/volume-one/babylon').read() main_text, footnotes_text = extract_main_content_and_footnotes(html) # 查看结果 print("=== 正文内容 ===") print(main_text) print("\n=== 脚注内容 ===") print(footnotes_text)
关键修改说明
去除重复链接:
- 新增排除
nav、footer、header标签,直接过滤页面导航栏、页脚区域的内容,避免重复链接混入正文。 - 先定位正文专属容器(
entry-content类的div),只提取该容器内的文本,彻底隔离其他区域的无关内容。
- 新增排除
分离脚注与正文:
单独定位脚注专属容器(footnotes类的div),独立提取该区域的文本,最终返回正文和脚注两个独立结果。优化文本干净度:
加入if t.strip()判断,过滤掉提取过程中产生的空字符串,让结果更整洁。
内容的提问来源于stack exchange,提问作者Talal Ghannam
相关产品推荐
相关产品推荐

