You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取新闻存入Pandas DataFrame:空对象致列表长度不一致问题

问题分析与解决方案

你的核心问题是部分页面未找到AMP链接时,没有向对应列表追加空值,导致AMP链接列表长度比其他列表短。原代码中对AMP链接的处理依赖findAll()循环,当页面无匹配元素时循环根本不会执行,自然不会追加空值。另外,原代码中标题、发布时间的逻辑也存在隐患——如果某页面没有对应元素,同样会导致列表长度不一致,只是这次刚好所有页面都有这些元素而已。

修正思路

对每个页面的四个字段,都采用「找单个元素→有值则取,无值则补空」的逻辑,确保每个页面都能向四个列表各追加一个元素,保证列表长度完全一致。

修正后的代码

import requests
from bs4 import BeautifulSoup
import re
import pandas as pd

def news_scraping_wienerzeitung():
    wienerzeitung_url='https://www.wienerzeitung.at/'
    html = requests.get(wienerzeitung_url)
    bsobj_3 = BeautifulSoup(html.content, 'html.parser')

    links = []
    for link in bsobj_3.find_all('a', attrs={'href': re.compile("^https://www.wienerzeitung.at/nachrichten")}):
        links.append(link['href'])

    lst_title = []
    lst_content = []
    lst_published = []
    lst_amp_link = []

    for l in links:
        page = requests.get(l)
        b = BeautifulSoup(page.content, 'html.parser')
        
        # 处理标题:存在则取文本,不存在补None
        title_elem = b.find('h1', {'class': 'article-title d-inline'})
        lst_title.append(title_elem.text.strip() if title_elem else None)
        
        # 处理内容:存在则取文本,不存在补空字符串
        content_elem = b.find('p', {'id': 'absatz1'})
        lst_content.append(content_elem.get_text().strip() if content_elem else '')
        
        # 处理发布时间:存在则取文本,不存在补None
        published_elem = b.find('span', {'class': 'article-published'})
        lst_published.append(published_elem.text.strip() if published_elem else None)
        
        # 处理AMP链接:存在则取href,不存在补None
        amp_elem = b.find('link', {'rel': 'amphtml'})
        lst_amp_link.append(amp_elem['href'] if amp_elem else None)

    # 验证各列表长度
    print(len(lst_published))
    print(len(lst_title))
    print(len(lst_content))
    print(len(lst_amp_link))
    
    # 转换为DataFrame
    df = pd.DataFrame({
        '标题': lst_title,
        '内容': lst_content,
        '发布时间': lst_published,
        'AMP链接': lst_amp_link
    })
    return df

# 执行并获取结果
result_df = news_scraping_wienerzeitung()

关键修改点

  1. 替换findAll()为find():每个页面的标题、发布时间等都是唯一元素,用find()更高效,也便于直接判断元素是否存在。
  2. 三元表达式处理空值:对每个元素做存在性判断,确保无论是否找到元素,都能向列表追加一个值(有效值或空值),保证所有列表长度一致。
  3. 语义化变量名:将lst1/lst2等改为lst_title/lst_content,提升代码可读性。
  4. 集成DataFrame生成:处理完列表后直接生成DataFrame,无需额外步骤。

内容的提问来源于stack exchange,提问作者Thulana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 09:53:22