Python中谷歌新闻指定日期范围链接爬取结果异常问题排查
问题:谷歌新闻日期筛选URL手动有效,但爬虫提取的链接不在指定日期范围
问题描述
我正在为小型机构开发Python媒体追踪器,通过构造谷歌新闻搜索URL并提取页面链接。手动打开构造好的URL能得到符合指定日期范围的结果,但代码爬取到的链接却不在该日期范围内。试过googlenews模块但不符合需求(谷歌新闻搜索功能与Google News模块逻辑不同)。
原代码:
import requests import re import bs4 import urllib import pandas as pd from bs4 import BeautifulSoup from bs4 import BeautifulSoup as bs from urllib.request import urlopen search = 'Center+for+Community+Alternatives' start_date = '5/2/2019' end_date = '5/10/2019' # The link to be searched initially link4 = 'https://www.google.com/search?q=' + search + '&tbs=cdr:1,cd_min:' + start_date + ',cd_max:' + end_date + '&tbm=nws' # Searching the link page = requests.get(link4) # Gathering the page content from the link soup = BeautifulSoup(page.content) # Finding all links links = soup.findAll("a") # Fixing up the links and putting them into a list empty = [] fixed_list = [] finished_list = [] for link in soup.find_all("a",href=re.compile("(?<=/url\?q=)(htt.*://.*)")): the = (re.split(":(?=http)",link["href"].replace("/url?q=",""))) empty.append(the) for link in empty: fixed_list.append(link[0]) for link in fixed_list: finished_list.append(link.split('&sa',1)[0]) # The list of links finished_list # The original link searched print(link4)
原因分析
- 反爬机制导致请求未获取筛选后内容:直接用
requests.get()发送请求时,未携带浏览器标识等关键请求头,谷歌会判定为非浏览器访问,返回未应用日期筛选的通用结果(或简化页面),而非手动访问时的筛选后内容。 - 链接提取逻辑冗余易出错:原代码用多层循环和复杂正则拆分链接,不仅效率低,还可能遗漏部分链接或处理错误。
修复方案
1. 添加浏览器请求头
在请求中加入模拟浏览器的User-Agent等头信息,让谷歌返回正确的筛选后页面。可通过浏览器开发者工具获取自己的真实User-Agent。
2. 优化链接提取逻辑
利用urllib.parse解析谷歌的跳转链接,替代复杂正则,更可靠地提取真实目标链接。
修改后的代码
import requests from bs4 import BeautifulSoup from urllib.parse import unquote, urlparse, parse_qs search = 'Center+for+Community+Alternatives' start_date = '5/2/2019' end_date = '5/10/2019' # 构造带日期筛选的谷歌新闻搜索URL link4 = f'https://www.google.com/search?q={search}&tbs=cdr:1,cd_min:{start_date},cd_max:{end_date}&tbm=nws' # 模拟浏览器请求头,替换为你自己的User-Agent headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } # 发送带请求头的请求 page = requests.get(link4, headers=headers) # 可取消注释查看返回页面内容,验证是否包含筛选后的新闻 # print(page.text) # 解析页面 soup = BeautifulSoup(page.content, 'html.parser') finished_list = [] # 提取所有谷歌跳转链接 for a_tag in soup.find_all('a', href=True): href = a_tag['href'] if href.startswith('/url?q='): # 解码并解析链接参数 decoded_href = unquote(href) parsed = urlparse(decoded_href) query_params = parse_qs(parsed.query) if 'q' in query_params: real_link = query_params['q'][0] # 去除额外参数 if '&sa=' in real_link: real_link = real_link.split('&sa=')[0] finished_list.append(real_link) # 去重避免重复链接 finished_list = list(set(finished_list)) # 输出结果 print("筛选后的新闻链接:") for link in finished_list: print(link) print("\n原搜索URL:") print(link4)
额外注意事项
- 若仍被拦截,可尝试添加
Cookie参数(从浏览器登录谷歌后的Cookie中提取关键值),但Cookie有有效期,需定期更新。 - 谷歌页面结构可能随时变化,若后续提取失效,需重新检查页面元素选择器。
内容的提问来源于stack exchange,提问作者Elijah Appelson
相关产品推荐
相关产品推荐

