You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取Inshorts网站时返回空列表的技术问题排查

爬取Inshorts返回空列表的原因及修复方案

问题分析

你的代码返回空列表主要有以下几个问题:

  • 错误使用find()而非find_all():soup.find()只会返回第一个匹配的元素,而非所有符合条件的元素集合。要获取所有新闻标题和内容容器,必须用find_all()。
  • find_all()无法直接调用.string:find_all()返回的是元素结果集(ResultSet),不是单个元素,不能直接通过.string提取文本。应该用find()定位单个目标元素后再获取文本。
  • 选择器逻辑偏差:原代码试图分别抓取所有标题容器和内容容器,这种方式容易出现元素数量不匹配的情况。更可靠的逻辑是先定位每个完整的新闻卡片,再从卡片内部提取标题和内容。

修复后的代码

import requests
from bs4 import BeautifulSoup

url = 'https://inshorts.com/en/read/technology'
news_data = []
news_category = url.split('/')[-1]

headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}
data = requests.get(url, headers=headers)

if data.status_code == 200:
    soup = BeautifulSoup(data.content, 'html.parser')
    # 获取所有新闻卡片
    news_cards = soup.find_all('div', class_='news-card')
    
    for card in news_cards:
        # 从卡片内提取标题
        headline = card.find('span', attrs={'itemprop': 'headline'}).string
        # 从卡片内提取内容
        article = card.find('div', attrs={'itemprop': 'articleBody'}).string
        news_data.append({
            'news_headline': headline,
            'news_article': article,
            'news_category': news_category
        })

print(news_data)

关键修复点说明

  • 改用find_all('div', class_='news-card')获取所有新闻卡片,确保每个卡片对应一条完整新闻,避免元素数量不匹配。
  • 在每个卡片内部使用find()定位标题和内容元素,直接获取单个元素的文本内容,避免结果集操作错误。
  • 移除了原代码中冗余的长度判断,因为每个新闻卡片内必然包含一个标题和一个内容块,逻辑更稳定。

内容的提问来源于stack exchange,提问作者Vighnesh Bhat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 10:25:01