Python谷歌新闻Web Scraping脚本无结果问题排查求助
解决Google News爬虫无结果的问题
你的爬虫没返回结果主要有两个核心问题:语法缩进错误,以及Google News页面结构已更新导致旧选择器失效。
1. 修复语法缩进错误
你的代码里有两处缩进问题,会直接导致代码运行报错或逻辑异常:
- 函数
scrapeGoogleNews内的RESULT = []没有缩进,属于函数外变量,函数内无法正确调用 - 示例循环里的所有
print语句未缩进,不属于循环块,无法遍历输出结果
2. 更新页面选择器
Google News的DOM结构已更新,你使用的旧类名(xrnccd、ipQwMb等)已不存在,导致soup.find_all返回空列表。以下是适配当前页面结构的修正代码:
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin def scrapeGoogleNews(): RESULT = [] base_url = "https://news.google.com" # 更新为较新的User-Agent,降低被拦截概率 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36" } response = requests.get(base_url, headers=headers) assert response.status_code == 200, f"请求失败,状态码:{response.status_code}" soup = BeautifulSoup(response.text, 'html.parser') # 选择当前Google News的文章容器标签 articles = soup.find_all("article") for article in articles: # 提取标题(跳过无标题条目) title_tag = article.find("h4") if not title_tag: continue title = title_tag.text.strip() # 提取描述(无描述则填充默认值) desc_tag = article.find("div", class_="QmrVtf") description = desc_tag.text.strip() if desc_tag else "无描述" # 提取并拼接新闻链接(处理相对路径) link_tag = article.find("a", class_="WwrzSb") if not link_tag: continue relative_url = link_tag.get("href").lstrip('.') news_url = urljoin(base_url, relative_url) RESULT.append([title, description, news_url]) return RESULT # 示例用法: news_articles = scrapeGoogleNews() for article in news_articles: print("Title:", article[0]) print("Description:", article[1]) print("URL:", article[2]) print()
额外注意事项
- Google News会频繁更新页面结构,后续若再次失效,需用浏览器F12开发者工具重新查看元素结构,更新选择器
- 频繁爬取可能触发反爬机制,建议添加请求间隔(如
time.sleep(1)),避免IP被封禁 - 确保网络环境可正常访问Google News
内容的提问来源于stack exchange,提问作者Dalerjon Nuriddinov
相关产品推荐
相关产品推荐

