使用Beautiful Soup抓取谷歌搜索结果无法提取描述文本怎么办?
问题成因
- 容器选择逻辑错误:当前代码中遍历的
container是搜索结果内的<a>标签,而搜索结果描述对应的<div>节点与该<a>标签为同级节点,同属于class="g"的父级<div>容器,在<a>标签内部检索描述节点必然返回空值。 - 方法语法混用:
select_one()方法仅接受CSS选择器字符串作为单参数,代码中传入第二个类名字典参数的写法是find()方法的语法,本身不符合select_one()的调用规则。 - 类名匹配策略不合理:谷歌搜索的节点类名会根据访问地区、UA标识、页面版本动态调整,硬编码全量类名匹配的容错率极低,很容易出现匹配失效的问题。
修复方案
调整容器遍历逻辑为直接匹配class="g"的父级搜索结果容器,修正选择器写法,同时增加节点存在性判断避免无描述内容时报错,修复后代码如下:
from bs4 import BeautifulSoup from urllib.parse import urljoin import requests base = "https://www.google.de" link = "https://www.google.de/search?q={}" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.150 Safari/537.36' } def grab_content(link): res = requests.get(link,headers=headers) soup = BeautifulSoup(res.text,"lxml") # 遍历父级搜索结果容器 for container in soup.select("div.g"): title_node = container.select_one("a[href^='http'][data-ved]:has(h3)") if not title_node: continue post_title = title_node.select_one("h3").get_text(strip=True) post_link = title_node.get('href') # 模糊匹配描述节点,避免类名变动失效 desc_node = container.select_one("div.VwiC3b, div.MUxGbd") post_description = desc_node.get_text(strip=True) if desc_node else "" yield post_title, post_link, post_description next_page = soup.select_one("a#pnnext") if next_page: next_page_link = urljoin(base, next_page.get("href")) yield from grab_content(next_page_link) if __name__ == '__main__': search_keyword = "python" qualified_link = link.format(search_keyword.replace(" ","+")) for item in grab_content(qualified_link): print(item)
运行效果
修复后输出结果如下,可正常获取描述内容:
('Welcome to Python.org', 'https://www.python.org/', 'The official home of the Python Programming Language.')
内容的提问来源于stack exchange,提问作者Raspberry Lemon
相关产品推荐
相关产品推荐

