You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas编写网页爬取脚本仅获最后一行数据的问题排查

问题分析与修复方案

核心问题点

  1. dropna操作未生效:Pandas的dropna方法不会修改原DataFrame,必须将返回结果重新赋值给df,否则空值行仍会留在数据中。
  2. 循环逻辑错误:字典d的构建和data.append(d)语句写在for循环外部,导致循环结束后仅将最后一次爬取的结果添加到列表,自然只能得到最后一行数据。
  3. 无效的CSV写入:循环前执行df.to_csv是将未处理的原始网站列表写入文件,完全没必要,应该在爬取完成后再写入结果。

修正后的代码

import pandas as pd
import requests
from bs4 import BeautifulSoup

# 读取CSV并过滤空值
df = pd.read_csv('~/Documents/websites.csv', usecols=['website'], delimiter=',')
df = df.dropna(subset=['website'])  # 关键:重新赋值保存过滤后的结果

data = []

for index, row in df.iterrows():
    website = row['website']
    try:
        # 增加异常处理,避免单个网站请求失败导致整个脚本中断
        response = requests.get(website)
        response.raise_for_status()  # 抛出HTTP错误
        content = response.text

        soup = BeautifulSoup(content, 'html.parser')
        
        tag1 = soup.find('span', class_='KeyStatisticsCard_field-text__GtuGd')
        tag2 = soup.find('div', class_='ant-col KeyStatisticsCard_field-info__gYdfV')
        tag3 = soup.find('a', class_='KeyStatisticsCard_ellipsis__TE9tk')
        tag4 = soup.find('span', class_='KeyStatisticsCard_ellipsis__TE9tk')

        # 构建字典并添加到列表,这一步必须放在循环内部
        d = {
            'tag1': tag1.text if tag1 else None,  # 增加判空,避免标签不存在时报错
            'tag2': tag2.text if tag2 else None,
            'tag3': tag3.text if tag3 else None,
            'tag4': tag4.text if tag4 else None
        }
        data.append(d)
    except Exception as e:
        # 捕获异常并记录,方便排查问题
        print(f"处理网站 {website} 时出错: {str(e)}")
        # 出错时也添加一条空数据或错误标记,保证数据行数匹配
        data.append({'tag1': None, 'tag2': None, 'tag3': None, 'tag4': f"Error: {str(e)}"})

# 转换为DataFrame并写入CSV
data_df = pd.DataFrame(data)
data_df.to_csv('/Users/Desktop/scraped_data.csv', header=True, index=False)

print(data_df)

关键修改说明

  • 将df = df.dropna(subset=['website'])赋值给原变量,确保空值行被过滤。
  • 把字典d的构建和data.append(d)移到for循环内部,每次爬取后立即保存结果。
  • 增加try-except异常处理,避免单个网站请求失败导致脚本中断,同时记录错误信息。
  • 给每个标签的text获取增加判空逻辑,防止标签不存在时抛出AttributeError。
  • 将CSV写入操作移到循环结束后,确保写入的是完整的爬取结果。

内容的提问来源于stack exchange,提问作者Zer0chance_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 18:32:22