爬取Inshorts网站时返回空列表的技术问题排查
爬取Inshorts返回空列表的原因及修复方案
问题分析
你的代码返回空列表主要有以下几个问题:
- 错误使用
find()而非find_all():soup.find()只会返回第一个匹配的元素,而非所有符合条件的元素集合。要获取所有新闻标题和内容容器,必须用find_all()。 find_all()无法直接调用.string:find_all()返回的是元素结果集(ResultSet),不是单个元素,不能直接通过.string提取文本。应该用find()定位单个目标元素后再获取文本。- 选择器逻辑偏差:原代码试图分别抓取所有标题容器和内容容器,这种方式容易出现元素数量不匹配的情况。更可靠的逻辑是先定位每个完整的新闻卡片,再从卡片内部提取标题和内容。
修复后的代码
import requests from bs4 import BeautifulSoup url = 'https://inshorts.com/en/read/technology' news_data = [] news_category = url.split('/')[-1] headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} data = requests.get(url, headers=headers) if data.status_code == 200: soup = BeautifulSoup(data.content, 'html.parser') # 获取所有新闻卡片 news_cards = soup.find_all('div', class_='news-card') for card in news_cards: # 从卡片内提取标题 headline = card.find('span', attrs={'itemprop': 'headline'}).string # 从卡片内提取内容 article = card.find('div', attrs={'itemprop': 'articleBody'}).string news_data.append({ 'news_headline': headline, 'news_article': article, 'news_category': news_category }) print(news_data)
关键修复点说明
- 改用
find_all('div', class_='news-card')获取所有新闻卡片,确保每个卡片对应一条完整新闻,避免元素数量不匹配。 - 在每个卡片内部使用
find()定位标题和内容元素,直接获取单个元素的文本内容,避免结果集操作错误。 - 移除了原代码中冗余的长度判断,因为每个新闻卡片内必然包含一个标题和一个内容块,逻辑更稳定。
内容的提问来源于stack exchange,提问作者Vighnesh Bhat
相关产品推荐
相关产品推荐

