使用Gnews包爬取Google News时标题不完整的解决方法
解决gnews爬取Google News标题被截断的问题
gnews返回的截断标题来源于Google搜索结果页的展示内容,要获取完整标题,需要通过新闻链接抓取原网页的标题信息,具体实现如下:
首先补充依赖库,用于请求原网页和解析内容:
pip install requests beautifulsoup4修改原有代码,添加抓取完整标题的逻辑:
from gnews import GNews import datetime import pandas as pd import requests from bs4 import BeautifulSoup # 配置请求头,模拟浏览器访问,避免被拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } start = datetime.date(2022, 1, 1) end = datetime.date(2022, 1, 4) google_news = GNews(start_date=start, end_date=end) rsp = google_news.get_news("IA|Inteligencia Artificial|inteligencia artificial|Inteligencia artificial") # 遍历每条新闻,获取完整标题 for item in rsp: news_url = item.get('url') if news_url: try: response = requests.get(news_url, headers=headers, timeout=10) response.encoding = response.apparent_encoding soup = BeautifulSoup(response.text, 'html.parser') # 提取原网页的title标签内容 full_title = soup.title.string.strip() if soup.title else '无法获取完整标题' # 替换原截断标题或新增完整标题字段 item['full_title'] = full_title except Exception as e: print(f"获取{news_url}标题失败: {str(e)}") item['full_title'] = '无法获取完整标题' # 打印包含完整标题的结果 for item in rsp: print(f"截断标题: {item['title']}") print(f"完整标题: {item['full_title']}") print("---")注意事项:
- 部分网站可能有反爬机制,若请求失败可调整
User-Agent或添加代理IP; - 少数网页的title标签可能包含冗余信息(比如网站名称后缀),可根据需求进一步处理字符串,去除无关内容。
- 部分网站可能有反爬机制,若请求失败可调整
内容的提问来源于stack exchange,提问作者Ana
相关产品推荐
相关产品推荐

