You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Gnews包爬取Google News时标题不完整的解决方法

解决gnews爬取Google News标题被截断的问题

gnews返回的截断标题来源于Google搜索结果页的展示内容,要获取完整标题,需要通过新闻链接抓取原网页的标题信息,具体实现如下:

  • 首先补充依赖库,用于请求原网页和解析内容:

    pip install requests beautifulsoup4
    
  • 修改原有代码,添加抓取完整标题的逻辑:

    from gnews import GNews
    import datetime
    import pandas as pd
    import requests
    from bs4 import BeautifulSoup
    
    # 配置请求头,模拟浏览器访问,避免被拦截
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    
    start = datetime.date(2022, 1, 1)
    end = datetime.date(2022, 1, 4)
    google_news = GNews(start_date=start, end_date=end)
    rsp = google_news.get_news("IA|Inteligencia Artificial|inteligencia artificial|Inteligencia artificial")
    
    # 遍历每条新闻,获取完整标题
    for item in rsp:
        news_url = item.get('url')
        if news_url:
            try:
                response = requests.get(news_url, headers=headers, timeout=10)
                response.encoding = response.apparent_encoding
                soup = BeautifulSoup(response.text, 'html.parser')
                # 提取原网页的title标签内容
                full_title = soup.title.string.strip() if soup.title else '无法获取完整标题'
                # 替换原截断标题或新增完整标题字段
                item['full_title'] = full_title
            except Exception as e:
                print(f"获取{news_url}标题失败: {str(e)}")
                item['full_title'] = '无法获取完整标题'
    
    # 打印包含完整标题的结果
    for item in rsp:
        print(f"截断标题: {item['title']}")
        print(f"完整标题: {item['full_title']}")
        print("---")
    
  • 注意事项:

    • 部分网站可能有反爬机制,若请求失败可调整User-Agent或添加代理IP;
    • 少数网页的title标签可能包含冗余信息(比如网站名称后缀),可根据需求进一步处理字符串,去除无关内容。

内容的提问来源于stack exchange,提问作者Ana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 03:16:05