如何解决DataFrame保存CSV报错及csv.DictReader对象异常问题
问题描述
- 爬虫逻辑运行正常,但保存数据到CSV时触发DataFrame相关报错
- 代码执行初期出现
<csv.DictReader object at 0x000001B5CCE8D240> []的异常提示 - 已卸载重装Pandas库,问题仍未解决
- 代码文件名为
multiple.py,爬虫代码如下:
import requests from bs4 import BeautifulSoup import time import pandas as pd def Gucci(url): headers={ 'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.5112.79 Safari/537.36'} r=requests.get(url,headers=headers) soup = BeautifulSoup(r.content, 'html.parser') content = soup.find_all('a',class_='product-tiles-grid-item-link js-ga-track') time.sleep(6) for items in content: Title=items.find('div',class_='product-tiles-grid-item-info').h2.text price=items.find('p',class_='price').span.text.replace("CHF", "") Store='GUCCI' Product_Type='Coat' Category= 'Mens' Country='Switzerland' #print(Title,price,Store,Product_Type,Category) skechers={ 'Product_Name':Title, 'Product_Price': price, 'Store':Store, 'Product_Type':Product_Type, 'Category': Category, 'Country':Country } skechersshoes.append(skechers) df = pd.DataFrame(skechersshoes) #print(df.head()) print('ALL DONE ') df.to_csv('Gucci.csv') print('Saved to csv file') return Gucci('https://www.gucci.com/ch/en_gb/ca/men/ready-to-wear-for-men/coats-for-men-c-men-readytowear-coats')
问题原因分析
- 核心问题:列表未初始化:代码中直接调用
skechersshoes.append(),但该列表从未定义,导致创建DataFrame时抛出异常 - 空数据风险:如果网页解析不到目标内容,
content为空会导致列表始终为空,同样会触发DataFrame相关报错 <csv.DictReader>提示:和当前爬虫逻辑无关,大概率是运行环境中残留的其他代码/变量导致
修复后的代码
import requests from bs4 import BeautifulSoup import time import pandas as pd def Gucci(url): # 初始化存储数据的列表 skechersshoes = [] headers={ 'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.5112.79 Safari/537.36' } r=requests.get(url,headers=headers) soup = BeautifulSoup(r.content, 'html.parser') content = soup.find_all('a',class_='product-tiles-grid-item-link js-ga-track') time.sleep(6) for items in content: # 捕获解析异常,避免单个商品解析失败中断程序 try: Title=items.find('div',class_='product-tiles-grid-item-info').h2.text.strip() price=items.find('p',class_='price').span.text.replace("CHF", "").strip() except AttributeError: continue Store='GUCCI' Product_Type='Coat' Category= 'Mens' Country='Switzerland' skechers={ 'Product_Name':Title, 'Product_Price': price, 'Store':Store, 'Product_Type':Product_Type, 'Category': Category, 'Country':Country } skechersshoes.append(skechers) # 检查是否抓取到数据,避免空DataFrame报错 if not skechersshoes: print('未抓取到任何数据') return df = pd.DataFrame(skechersshoes) print('ALL DONE ') df.to_csv('Gucci.csv', index=False) # 移除多余索引列 print('Saved to csv file') Gucci('https://www.gucci.com/ch/en_gb/ca/men/ready-to-wear-for-men/coats-for-men-c-men-readytowear-coats')
修复说明
- 在函数开头初始化
skechersshoes = [],解决列表未定义的核心问题 - 添加
try-except捕获解析异常,避免单个商品解析失败导致程序终止 - 增加空列表检查,防止因无数据创建空DataFrame引发报错
- 保存CSV时添加
index=False,避免生成不必要的索引列 - 调用
.strip()去除文本多余空格,优化数据整洁度
内容的提问来源于stack exchange,提问作者Muhammad Umer
相关产品推荐
相关产品推荐

