You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup从谷歌搜索结果提取URL时始终得到空列表

解决谷歌搜索结果URL提取为空的问题

你的代码失效核心原因有两个:

  • 谷歌搜索结果的HTML结构已更新,你依赖的class="r"容器不再用于包裹搜索结果
  • 未处理URL编码和分页逻辑,无法正确获取多页结果

修正后的代码实现

import requests
from bs4 import BeautifulSoup
from urllib.parse import quote, urlparse, parse_qs

def get_location_info(location, max_results=100):
    query = quote(f"{location} information")
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept-Language': 'en-US,en;q=0.5',
        'Referer': 'https://www.google.com/'
    }
    
    websites = []
    page = 0
    
    while len(websites) < max_results:
        # 构造带分页参数的URL,谷歌默认每页返回10条结果
        url = f"https://www.google.com/search?q={query}&start={page*10}"
        response = requests.get(url, headers=headers)
        
        if response.status_code != 200:
            print("请求被谷歌拦截,请检查请求头或尝试添加浏览器Cookie")
            break
            
        soup = BeautifulSoup(response.text, 'html.parser')
        # 定位当前页的搜索结果容器(谷歌当前结构为每个结果包裹在div.g中)
        result_containers = soup.find_all("div", class_="g")
        
        if not result_containers:
            print("当前页无结果,已到达搜索末尾")
            break
            
        for container in result_containers:
            link_tag = container.find("a", href=True)
            if link_tag and link_tag['href'].startswith('/url?q='):
                # 解析谷歌重定向链接中的真实URL
                parsed_url = urlparse(link_tag['href'])
                real_url = parse_qs(parsed_url.query).get('q', [None])[0]
                if real_url and real_url not in websites:
                    websites.append(real_url)
                    if len(websites) >= max_results:
                        break
        
        page += 1
    
    return websites

# 测试调用
location = "Athens"
print(get_location_info(location))

关键改动说明

  1. URL编码:用urllib.parse.quote处理查询字符串,避免特殊字符导致的请求错误
  2. 更新选择器:改用当前谷歌搜索结果的容器div.g,提取/url?q=格式的链接并解析真实地址
  3. 分页逻辑:通过start参数循环请求多页,直到收集到100个有效链接
  4. 增强请求头:使用现代浏览器的User-Agent,添加Accept-Language和Referer降低反爬拦截概率

额外注意事项

  • 如果仍被拦截,可从浏览器开发者工具中复制真实Cookie添加到请求头
  • 谷歌页面结构可能随时更新,若后续失效需重新检查元素选择器

内容的提问来源于stack exchange,提问作者Priniotis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 21:20:33