使用Pandas编写网页爬取脚本仅获最后一行数据的问题排查
问题分析与修复方案
核心问题点
dropna操作未生效:Pandas的dropna方法不会修改原DataFrame,必须将返回结果重新赋值给df,否则空值行仍会留在数据中。- 循环逻辑错误:字典
d的构建和data.append(d)语句写在for循环外部,导致循环结束后仅将最后一次爬取的结果添加到列表,自然只能得到最后一行数据。 - 无效的CSV写入:循环前执行
df.to_csv是将未处理的原始网站列表写入文件,完全没必要,应该在爬取完成后再写入结果。
修正后的代码
import pandas as pd import requests from bs4 import BeautifulSoup # 读取CSV并过滤空值 df = pd.read_csv('~/Documents/websites.csv', usecols=['website'], delimiter=',') df = df.dropna(subset=['website']) # 关键:重新赋值保存过滤后的结果 data = [] for index, row in df.iterrows(): website = row['website'] try: # 增加异常处理,避免单个网站请求失败导致整个脚本中断 response = requests.get(website) response.raise_for_status() # 抛出HTTP错误 content = response.text soup = BeautifulSoup(content, 'html.parser') tag1 = soup.find('span', class_='KeyStatisticsCard_field-text__GtuGd') tag2 = soup.find('div', class_='ant-col KeyStatisticsCard_field-info__gYdfV') tag3 = soup.find('a', class_='KeyStatisticsCard_ellipsis__TE9tk') tag4 = soup.find('span', class_='KeyStatisticsCard_ellipsis__TE9tk') # 构建字典并添加到列表,这一步必须放在循环内部 d = { 'tag1': tag1.text if tag1 else None, # 增加判空,避免标签不存在时报错 'tag2': tag2.text if tag2 else None, 'tag3': tag3.text if tag3 else None, 'tag4': tag4.text if tag4 else None } data.append(d) except Exception as e: # 捕获异常并记录,方便排查问题 print(f"处理网站 {website} 时出错: {str(e)}") # 出错时也添加一条空数据或错误标记,保证数据行数匹配 data.append({'tag1': None, 'tag2': None, 'tag3': None, 'tag4': f"Error: {str(e)}"}) # 转换为DataFrame并写入CSV data_df = pd.DataFrame(data) data_df.to_csv('/Users/Desktop/scraped_data.csv', header=True, index=False) print(data_df)
关键修改说明
- 将
df = df.dropna(subset=['website'])赋值给原变量,确保空值行被过滤。 - 把字典
d的构建和data.append(d)移到for循环内部,每次爬取后立即保存结果。 - 增加
try-except异常处理,避免单个网站请求失败导致脚本中断,同时记录错误信息。 - 给每个标签的
text获取增加判空逻辑,防止标签不存在时抛出AttributeError。 - 将CSV写入操作移到循环结束后,确保写入的是完整的爬取结果。
内容的提问来源于stack exchange,提问作者Zer0chance_
相关产品推荐
相关产品推荐

