如何用Beautiful Soup获取可见文本?能否检测各类隐藏元素?
问题
我查阅了Stack Overflow上的《BeautifulSoup Grab Visible Webpage Text》一文,发现其中的公认解决方案存在将隐藏文本误判为可见文本的问题。
以下是用XPATH筛选非隐藏复选框的示例:
CHECK_BOX_XPATH = "//input[(@type='checkbox')" " and(not(@style='display: none;')) and(not(@visibility='hidden')) and (not(@hidden)) and" " (not(@disabled)) and (not(contains(@class,'disabled')))]"
基于包含完整HTML及CSS属性的源码,请问Beautiful Soup能否检测上述示例中的这些隐藏元素情况,避免返回其文本?
附该问题的公认解决方案代码:
from bs4 import BeautifulSoup from bs4.element import Comment import urllib.request def tag_visible(element): if element.parent.name in ['style', 'script', 'head', 'title', 'meta', '[document]']: return False if isinstance(element, Comment): return False return True def text_from_html(body): soup = BeautifulSoup(body, 'html.parser') texts = soup.findAll(text=True) visible_texts = filter(tag_visible, texts) return u" ".join(t.strip() for t in visible_texts) html = urllib.request.urlopen('http://www.nytimes.com/2009/12/21/us/21storm.html').read() print(text_from_html(html))
回答
Beautiful Soup本身不具备直接判断元素视觉可见性的能力,它只是一个HTML/XML解析库,只能处理HTML结构和属性,无法解析CSS样式的实际渲染效果,也模拟不了浏览器的渲染逻辑。但我们可以通过扩展原方案中的tag_visible函数,手动检查元素的相关属性,来过滤掉你示例中提到的隐藏元素。
针对你给出的XPATH筛选条件,我们可以修改tag_visible函数,增加对元素及父元素的属性检查:
from bs4 import BeautifulSoup from bs4.element import Comment import urllib.request def tag_visible(element): # 过滤基础不可见标签 if element.parent.name in ['style', 'script', 'head', 'title', 'meta', '[document]']: return False if isinstance(element, Comment): return False # 递归检查当前元素及其所有父元素的隐藏相关属性 current = element while current is not None: # 检查hidden属性 if current.has_attr('hidden'): return False # 检查disabled属性或含disabled的class if current.has_attr('disabled') or (current.has_attr('class') and 'disabled' in current.get('class')): return False # 检查内联样式中的隐藏规则 if current.has_attr('style'): style_content = current.get('style').lower() if 'display: none' in style_content or 'visibility: hidden' in style_content: return False current = current.parent return True def text_from_html(body): soup = BeautifulSoup(body, 'html.parser') texts = soup.findAll(text=True) visible_texts = filter(tag_visible, texts) return " ".join(t.strip() for t in visible_texts) # 测试调用 html = urllib.request.urlopen('http://www.nytimes.com/2009/12/21/us/21storm.html').read() print(text_from_html(html))
需要注意的是,这种方式存在局限性:
- 只能覆盖内联样式和直接属性的情况,无法处理外部CSS文件或
<style>标签中定义的隐藏规则(比如.hidden { display: none; }) - 无法识别通过JavaScript动态修改的元素可见性
- 无法判断元素是否被其他元素遮挡、超出视口等视觉层面的不可见情况
如果需要完全准确地获取页面可见文本,建议结合Selenium这类能模拟浏览器渲染的工具,先让浏览器完成页面渲染,再提取可见内容。
内容的提问来源于stack exchange,提问作者Ariel
相关产品推荐
相关产品推荐

