Python多页网页爬取数据存入CSV仅保留最后列表的技术问题
问题原因
你遇到的问题核心是每次循环都覆盖了之前的数据:
- 每次进入内层循环都会重新创建
list1/list2/list3/list4,之前收集的数据被清空 - 每次循环都新建DataFrame并调用
to_csv(),默认写入模式为覆盖(mode='w'),所以最终只有最后一次循环的数据被保留
解决步骤
- 提前初始化全局数据容器:在所有循环外面创建空列表,用来收集每一条爬取到的完整数据
- 循环内只收集数据:每次爬取到
list4后,将其添加到全局容器中,不要每次都写CSV - 统一写入CSV:所有数据爬取完成后,再把全局容器转换成DataFrame,处理长度不一致的问题后写入CSV
修改后的代码
# 初始化全局数据容器,放在所有循环外面 all_data = [] # 排除的URL用集合存储,判断更高效 exclude_urls = { 'https://www.intel.com/content/www/us/en/products/sku/201889/intel-core-i310325-processor-8m-cache-up-to-4-70-ghz/specifications.html', 'https://www.intel.com/content/www/us/en/products/sku/197123/intel-core-i31000g4-processor-4m-cache-up-to-3-20-ghz/specifications.html', 'https://www.intel.com/content/www/us/en/products/sku/97930/intel-atom-processor-c3508-8m-cache-up-to-1-60-ghz/specifications.html' } for product in products: prod = 'https://www.intel.com' + product['href'] html_text4 = requests.get(prod).text soup4 = BeautifulSoup(html_text4, 'lxml') processors3 = soup4.find_all('div', {'class': 'add-compare-wrap'}) for processor3 in processors3: proc3 = 'https://www.intel.com' + processor3.a['href'] if proc3 not in exclude_urls: html_text5 = requests.get(proc3).text soup5 = BeautifulSoup(html_text5, 'lxml') essentials = soup5.find('div', {'id': 'specs-1-0-0'}).find_all('div', {'class': 'row tech-section-row'}) cpu_specifications = soup5.find('div', {'id': 'specs-1-0-1'}).find_all('div', {'class': 'row tech-section-row'}) package = soup5.find_all('div', {'class': 'tech-section'}) list1 = [] list2 = [] list3 = [] for ess in essentials: essential = ess.text list1.append(essential) for cpu in cpu_specifications: cpu_specification = cpu.text list2.append(cpu_specification) for p in package: p2 = p.find_all('h3') x = 'Package Specifications' for p3 in p2: p4 = p3.text if p4 == x: p3 = p.find_all('div', {'class': 'row tech-section-row'}) for package_specifications in p3: package_specification = package_specifications.text list3.append(package_specification) list4 = list1 + list2 + list3 # 将单条数据添加到全局容器,替换原有的单次写入操作 all_data.append(list4) # 所有数据爬取完成后,统一处理并写入CSV # 处理长度不一致问题:自动补全缺失值为空字符串 df = pd.DataFrame(all_data).fillna('') df.to_csv('file.csv', header=False, index=False)
额外优化建议
- 添加请求头:比如
requests.get(prod, headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'}),避免被网站反爬拦截 - 加入延时:每次请求后添加
time.sleep(1),减轻服务器压力,降低被封IP的风险 - 异常处理:对
requests.get()和soup.find()等操作添加try-except块,避免单条数据爬取失败导致整个程序中断
内容的提问来源于stack exchange,提问作者AK0901
相关产品推荐
相关产品推荐

