Python爬虫写入CSV重复覆盖数据,如何实现追加写入?
解决CSV追加写入问题的方案
你的问题根源有两个:
- 每次循环用
'w'模式打开文件,会清空原有内容并重新写入,导致之前页面的数据被覆盖 - 表头在每次循环中都被重复写入,即使改成追加模式也会出现多次表头的问题
修改方案1:一次性打开文件,全程写入(推荐)
把文件打开操作放到循环外面,先写入表头,再循环爬取每页数据并写入,这样既避免重复打开文件,也不会覆盖数据:
import requests from bs4 import BeautifulSoup from csv import writer # 打开文件,用'w'模式先写入表头 with open('Flipkart.csv', 'w', encoding='utf8', newline='') as f: thewriter = writer(f) header = ('Title', 'Specification', 'price', 'Rating Out of 5') thewriter.writerow(header) # 循环爬取每页数据 for page in range(1, 10): url = 'https://www.flipkart.com/search?q=laptops&otracker=search&otracker1=search&marketplace=FLIPKART&as-show=on&as=off&page={page}'.format(page=page) req = requests.get(url) soup = BeautifulSoup(req.content, 'html.parser') links = soup.find_all('div', class_='_2kHMtA') for link in links: title = link.find('div', class_='_4rR01T').text Specification = link.find('ul', class_='_1xgFaf').text price = link.find('div', class_='_30jeq3 _1_WHN1').text Rating = link.find('span', class_='_1lRcqv') Rating = Rating.text if Rating else 'N/A' info = [title, Specification, price, Rating] thewriter.writerow(info)
修改方案2:分情况处理表头和追加
如果必须在循环内打开文件,可以判断是否为第一页,第一页用'w'模式写入表头,后续页面用'a'模式追加数据:
import requests from bs4 import BeautifulSoup from csv import writer for page in range(1, 10): url = 'https://www.flipkart.com/search?q=laptops&otracker=search&otracker1=search&marketplace=FLIPKART&as-show=on&as=off&page={page}'.format(page=page) req = requests.get(url) soup = BeautifulSoup(req.content, 'html.parser') links = soup.find_all('div', class_='_2kHMtA') # 判断是否为第一页,决定打开模式 if page == 1: mode = 'w' write_header = True else: mode = 'a' write_header = False with open('Flipkart.csv', mode, encoding='utf8', newline='') as f: thewriter = writer(f) if write_header: header = ('Title', 'Specification', 'price', 'Rating Out of 5') thewriter.writerow(header) for link in links: title = link.find('div', class_='_4rR01T').text Specification = link.find('ul', class_='_1xgFaf').text price = link.find('div', class_='_30jeq3 _1_WHN1').text Rating = link.find('span', class_='_1lRcqv') Rating = Rating.text if Rating else 'N/A' info = [title, Specification, price, Rating] thewriter.writerow(info)
关键修改点说明
- 替换文件打开模式:
'w'改为'a'(追加模式),但要注意表头只写一次 - 优化文件操作:方案1只打开一次文件,比方案2多次打开关闭更高效
- 简化Rating判断:用三元表达式替代if-else,代码更简洁
内容的提问来源于stack exchange,提问作者Syed
相关产品推荐
相关产品推荐

