网页抓取新闻存入Pandas DataFrame:空对象致列表长度不一致问题
问题分析与解决方案
你的核心问题是部分页面未找到AMP链接时,没有向对应列表追加空值,导致AMP链接列表长度比其他列表短。原代码中对AMP链接的处理依赖findAll()循环,当页面无匹配元素时循环根本不会执行,自然不会追加空值。另外,原代码中标题、发布时间的逻辑也存在隐患——如果某页面没有对应元素,同样会导致列表长度不一致,只是这次刚好所有页面都有这些元素而已。
修正思路
对每个页面的四个字段,都采用「找单个元素→有值则取,无值则补空」的逻辑,确保每个页面都能向四个列表各追加一个元素,保证列表长度完全一致。
修正后的代码
import requests from bs4 import BeautifulSoup import re import pandas as pd def news_scraping_wienerzeitung(): wienerzeitung_url='https://www.wienerzeitung.at/' html = requests.get(wienerzeitung_url) bsobj_3 = BeautifulSoup(html.content, 'html.parser') links = [] for link in bsobj_3.find_all('a', attrs={'href': re.compile("^https://www.wienerzeitung.at/nachrichten")}): links.append(link['href']) lst_title = [] lst_content = [] lst_published = [] lst_amp_link = [] for l in links: page = requests.get(l) b = BeautifulSoup(page.content, 'html.parser') # 处理标题:存在则取文本,不存在补None title_elem = b.find('h1', {'class': 'article-title d-inline'}) lst_title.append(title_elem.text.strip() if title_elem else None) # 处理内容:存在则取文本,不存在补空字符串 content_elem = b.find('p', {'id': 'absatz1'}) lst_content.append(content_elem.get_text().strip() if content_elem else '') # 处理发布时间:存在则取文本,不存在补None published_elem = b.find('span', {'class': 'article-published'}) lst_published.append(published_elem.text.strip() if published_elem else None) # 处理AMP链接:存在则取href,不存在补None amp_elem = b.find('link', {'rel': 'amphtml'}) lst_amp_link.append(amp_elem['href'] if amp_elem else None) # 验证各列表长度 print(len(lst_published)) print(len(lst_title)) print(len(lst_content)) print(len(lst_amp_link)) # 转换为DataFrame df = pd.DataFrame({ '标题': lst_title, '内容': lst_content, '发布时间': lst_published, 'AMP链接': lst_amp_link }) return df # 执行并获取结果 result_df = news_scraping_wienerzeitung()
关键修改点
- 替换
findAll()为find():每个页面的标题、发布时间等都是唯一元素,用find()更高效,也便于直接判断元素是否存在。 - 三元表达式处理空值:对每个元素做存在性判断,确保无论是否找到元素,都能向列表追加一个值(有效值或空值),保证所有列表长度一致。
- 语义化变量名:将lst1/lst2等改为lst_title/lst_content,提升代码可读性。
- 集成DataFrame生成:处理完列表后直接生成DataFrame,无需额外步骤。
内容的提问来源于stack exchange,提问作者Thulana
相关产品推荐
相关产品推荐

