Python使用pandas输出CSV文件时如何避免出现重复内容
CSV重复问题解决方法
重复原因
- 你定义的
words和meanings是全局累加的列表,每次查询新词都会把历史所有查询记录整合成表,再用追加模式写入CSV,相当于每查一个新词就会把之前所有记录重复写入一次 - 没有做重复输入校验,多次输入同一个词汇也会多次写入
修复方案
代码逻辑做了针对初学者的简化,非常好理解:
- 程序启动时先读取已有CSV里的所有词汇,存到查询速度极快的集合里做去重校验
- 每次查询只写入当前这一条词汇的结果,不会连带历史记录重复写入
- 自动处理首次运行无CSV文件的场景,无需手动创建文件
修改后代码
import requests from bs4 import BeautifulSoup import pandas as pd import os # 固定CSV存储路径 filepath = 'C:/Users/dict1.csv' # 存储已经爬过的词汇,用来快速判重 saved_words = set() # 检查CSV文件是否已存在 if os.path.exists(filepath): # 读取已有内容,把存过的词放到集合里 old_data = pd.read_csv(filepath) saved_words = set(old_data['words'].tolist()) else: # 文件不存在就先创建,写入表头 pd.DataFrame(columns=['words', 'meanings']).to_csv(filepath, index=False) while True: spell = input("spell: ") # 校验词汇是否已经存过 if spell in saved_words: print(f"{spell} 已经存储过,无需重复爬取") continue # 原爬取逻辑保留 r = requests.get("http://www.urbandictionary.com/define.php?term={}".format(spell)) r.encoding = r.apparent_encoding soup = BeautifulSoup(r.content, features="lxml") meaning = soup.find("div", attrs={"class": "meaning"}).get_text() print(meaning) # 仅将当前查询的单条数据写入CSV new_row = pd.DataFrame({ 'words': [spell], 'meanings': [meaning] }) new_row.to_csv(filepath, mode='a', index=False, header=False) # 把当前词汇加入已存集合,避免本次运行期间重复输入 saved_words.add(spell)
额外提示
如果遇到爬取失败的情况,可以加个try-except捕获异常避免程序直接崩溃,你初学的话可以先把上面的逻辑跑通再逐步加功能。
内容的提问来源于stack exchange,提问作者Takahiro
相关产品推荐
相关产品推荐

