You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中谷歌新闻指定日期范围链接爬取结果异常问题排查

问题:谷歌新闻日期筛选URL手动有效,但爬虫提取的链接不在指定日期范围

问题描述

我正在为小型机构开发Python媒体追踪器,通过构造谷歌新闻搜索URL并提取页面链接。手动打开构造好的URL能得到符合指定日期范围的结果,但代码爬取到的链接却不在该日期范围内。试过googlenews模块但不符合需求(谷歌新闻搜索功能与Google News模块逻辑不同)。

原代码:

import requests
import re
import bs4
import urllib
import pandas as pd
from bs4 import BeautifulSoup
from bs4 import BeautifulSoup as bs
from urllib.request import urlopen

search = 'Center+for+Community+Alternatives'
start_date = '5/2/2019'
end_date = '5/10/2019'

# The link to be searched initially 
link4 = 'https://www.google.com/search?q=' + search + '&tbs=cdr:1,cd_min:' + start_date + ',cd_max:' + end_date + '&tbm=nws'

# Searching the link
page = requests.get(link4)

# Gathering the page content from the link
soup = BeautifulSoup(page.content)

# Finding all links
links = soup.findAll("a")

# Fixing up the links and putting them into a list
empty = []
fixed_list = []
finished_list = []
for link in  soup.find_all("a",href=re.compile("(?<=/url\?q=)(htt.*://.*)")):
    the = (re.split(":(?=http)",link["href"].replace("/url?q=","")))
    empty.append(the)
for link in empty:
    fixed_list.append(link[0])
for link in fixed_list:
    finished_list.append(link.split('&sa',1)[0])


# The list of links
finished_list

# The original link searched
print(link4)

原因分析

  • 反爬机制导致请求未获取筛选后内容:直接用requests.get()发送请求时,未携带浏览器标识等关键请求头,谷歌会判定为非浏览器访问,返回未应用日期筛选的通用结果(或简化页面),而非手动访问时的筛选后内容。
  • 链接提取逻辑冗余易出错:原代码用多层循环和复杂正则拆分链接,不仅效率低,还可能遗漏部分链接或处理错误。

修复方案

1. 添加浏览器请求头

在请求中加入模拟浏览器的User-Agent等头信息,让谷歌返回正确的筛选后页面。可通过浏览器开发者工具获取自己的真实User-Agent。

2. 优化链接提取逻辑

利用urllib.parse解析谷歌的跳转链接,替代复杂正则,更可靠地提取真实目标链接。

修改后的代码

import requests
from bs4 import BeautifulSoup
from urllib.parse import unquote, urlparse, parse_qs

search = 'Center+for+Community+Alternatives'
start_date = '5/2/2019'
end_date = '5/10/2019'

# 构造带日期筛选的谷歌新闻搜索URL
link4 = f'https://www.google.com/search?q={search}&tbs=cdr:1,cd_min:{start_date},cd_max:{end_date}&tbm=nws'

# 模拟浏览器请求头,替换为你自己的User-Agent
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

# 发送带请求头的请求
page = requests.get(link4, headers=headers)
# 可取消注释查看返回页面内容,验证是否包含筛选后的新闻
# print(page.text)

# 解析页面
soup = BeautifulSoup(page.content, 'html.parser')

finished_list = []
# 提取所有谷歌跳转链接
for a_tag in soup.find_all('a', href=True):
    href = a_tag['href']
    if href.startswith('/url?q='):
        # 解码并解析链接参数
        decoded_href = unquote(href)
        parsed = urlparse(decoded_href)
        query_params = parse_qs(parsed.query)
        if 'q' in query_params:
            real_link = query_params['q'][0]
            # 去除额外参数
            if '&sa=' in real_link:
                real_link = real_link.split('&sa=')[0]
            finished_list.append(real_link)

# 去重避免重复链接
finished_list = list(set(finished_list))

# 输出结果
print("筛选后的新闻链接:")
for link in finished_list:
    print(link)
print("\n原搜索URL:")
print(link4)

额外注意事项

  • 若仍被拦截,可尝试添加Cookie参数(从浏览器登录谷歌后的Cookie中提取关键值),但Cookie有有效期,需定期更新。
  • 谷歌页面结构可能随时变化,若后续提取失效,需重新检查页面元素选择器。

内容的提问来源于stack exchange,提问作者Elijah Appelson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 22:15:39