使用BeautifulSoup从谷歌搜索结果提取URL时始终得到空列表
解决谷歌搜索结果URL提取为空的问题
你的代码失效核心原因有两个:
- 谷歌搜索结果的HTML结构已更新,你依赖的
class="r"容器不再用于包裹搜索结果 - 未处理URL编码和分页逻辑,无法正确获取多页结果
修正后的代码实现
import requests from bs4 import BeautifulSoup from urllib.parse import quote, urlparse, parse_qs def get_location_info(location, max_results=100): query = quote(f"{location} information") headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.google.com/' } websites = [] page = 0 while len(websites) < max_results: # 构造带分页参数的URL,谷歌默认每页返回10条结果 url = f"https://www.google.com/search?q={query}&start={page*10}" response = requests.get(url, headers=headers) if response.status_code != 200: print("请求被谷歌拦截,请检查请求头或尝试添加浏览器Cookie") break soup = BeautifulSoup(response.text, 'html.parser') # 定位当前页的搜索结果容器(谷歌当前结构为每个结果包裹在div.g中) result_containers = soup.find_all("div", class_="g") if not result_containers: print("当前页无结果,已到达搜索末尾") break for container in result_containers: link_tag = container.find("a", href=True) if link_tag and link_tag['href'].startswith('/url?q='): # 解析谷歌重定向链接中的真实URL parsed_url = urlparse(link_tag['href']) real_url = parse_qs(parsed_url.query).get('q', [None])[0] if real_url and real_url not in websites: websites.append(real_url) if len(websites) >= max_results: break page += 1 return websites # 测试调用 location = "Athens" print(get_location_info(location))
关键改动说明
- URL编码:用
urllib.parse.quote处理查询字符串,避免特殊字符导致的请求错误 - 更新选择器:改用当前谷歌搜索结果的容器
div.g,提取/url?q=格式的链接并解析真实地址 - 分页逻辑:通过
start参数循环请求多页,直到收集到100个有效链接 - 增强请求头:使用现代浏览器的User-Agent,添加
Accept-Language和Referer降低反爬拦截概率
额外注意事项
- 如果仍被拦截,可从浏览器开发者工具中复制真实Cookie添加到请求头
- 谷歌页面结构可能随时更新,若后续失效需重新检查元素选择器
内容的提问来源于stack exchange,提问作者Priniotis
相关产品推荐
相关产品推荐

