You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Bs4返回[]如何修复?Google无结果页面文本识别问题

修复Google无搜索结果页的文本识别问题

问题原因

  • 爬虫拦截:直接用requests.get请求Google搜索会被识别为非浏览器请求,返回人机验证页面而非真实结果页,导致找不到目标文本。
  • 精确匹配失效:soup.find_all(text="Try different keywords")是精确匹配,但Google页面的提示文本可能包含空格、换行或嵌套在多标签中,精确匹配无法命中。
  • 提示文本更新:Google无结果时的提示文本可能已变更,当前常见表述包括"No results found"、"did not match any documents"等。

修复后代码

import requests
from bs4 import BeautifulSoup

# 模拟浏览器请求头,规避爬虫拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

site = input("Input the domain of the top level domain: ")
textsearch = input("Input what text you're looking for: ")

# 构造搜索URL
url = f"https://www.google.com/search?q=site%3A.{site}+%22{textsearch}%22"

# 带请求头发送请求,同时检查请求是否成功
r = requests.get(url, headers=headers)
r.raise_for_status()

soup = BeautifulSoup(r.text, 'html.parser')

# 定义无结果时的常见提示文本,兼容页面更新
no_result_texts = ["Try different keywords", "No results found", "did not match any documents"]
is_empty_page = False

# 遍历检查所有可能的提示文本,用模糊匹配替代精确匹配
for target_text in no_result_texts:
    # 排除脚本、样式标签,仅检查可见文本,同时去除首尾空白
    match = soup.find_all(lambda tag: tag.name not in ['script', 'style'] and target_text in tag.get_text(strip=True))
    if match:
        is_empty_page = True
        break

# 输出判断结果
print("搜索返回空白/无结果页面" if is_empty_page else "搜索有结果")

关键修复说明

  • 添加请求头:模拟真实浏览器请求,避免Google的爬虫拦截机制。
  • 模糊匹配文本:用lambda函数实现包含匹配,同时去除文本首尾空白,解决精确匹配失效问题。
  • 多文本兼容:同时检查多个无结果提示文本,避免因Google页面更新导致识别失败。

内容的提问来源于stack exchange,提问作者user14809494

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 00:30:16