You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多页网页爬取数据存入CSV仅保留最后列表的技术问题

问题原因

你遇到的问题核心是每次循环都覆盖了之前的数据:

  • 每次进入内层循环都会重新创建list1/list2/list3/list4,之前收集的数据被清空
  • 每次循环都新建DataFrame并调用to_csv(),默认写入模式为覆盖(mode='w'),所以最终只有最后一次循环的数据被保留
解决步骤
  1. 提前初始化全局数据容器:在所有循环外面创建空列表,用来收集每一条爬取到的完整数据
  2. 循环内只收集数据:每次爬取到list4后,将其添加到全局容器中,不要每次都写CSV
  3. 统一写入CSV:所有数据爬取完成后,再把全局容器转换成DataFrame,处理长度不一致的问题后写入CSV
修改后的代码
# 初始化全局数据容器,放在所有循环外面
all_data = []
# 排除的URL用集合存储,判断更高效
exclude_urls = {
    'https://www.intel.com/content/www/us/en/products/sku/201889/intel-core-i310325-processor-8m-cache-up-to-4-70-ghz/specifications.html',
    'https://www.intel.com/content/www/us/en/products/sku/197123/intel-core-i31000g4-processor-4m-cache-up-to-3-20-ghz/specifications.html',
    'https://www.intel.com/content/www/us/en/products/sku/97930/intel-atom-processor-c3508-8m-cache-up-to-1-60-ghz/specifications.html'
}

for product in products:
    prod = 'https://www.intel.com' + product['href']
    html_text4 = requests.get(prod).text
    soup4 = BeautifulSoup(html_text4, 'lxml')
    processors3 = soup4.find_all('div', {'class': 'add-compare-wrap'})
    for processor3 in processors3:
        proc3 = 'https://www.intel.com' + processor3.a['href']
        if proc3 not in exclude_urls:
            html_text5 = requests.get(proc3).text
            soup5 = BeautifulSoup(html_text5, 'lxml')
            essentials = soup5.find('div', {'id': 'specs-1-0-0'}).find_all('div', {'class': 'row tech-section-row'})
            cpu_specifications = soup5.find('div', {'id': 'specs-1-0-1'}).find_all('div', {'class': 'row tech-section-row'})
            package = soup5.find_all('div', {'class': 'tech-section'})
            list1 = []
            list2 = []
            list3 = []
            for ess in essentials:
                essential = ess.text
                list1.append(essential)
            for cpu in cpu_specifications:
                cpu_specification = cpu.text
                list2.append(cpu_specification)
            for p in package:
                p2 = p.find_all('h3')
                x = 'Package Specifications'
                for p3 in p2:
                    p4 = p3.text
                if p4 == x:
                    p3 = p.find_all('div', {'class': 'row tech-section-row'})
                    for package_specifications in p3:
                        package_specification = package_specifications.text
                        list3.append(package_specification)
            list4 = list1 + list2 + list3
            # 将单条数据添加到全局容器,替换原有的单次写入操作
            all_data.append(list4)

# 所有数据爬取完成后,统一处理并写入CSV
# 处理长度不一致问题:自动补全缺失值为空字符串
df = pd.DataFrame(all_data).fillna('')
df.to_csv('file.csv', header=False, index=False)
额外优化建议
  • 添加请求头:比如requests.get(prod, headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'}),避免被网站反爬拦截
  • 加入延时:每次请求后添加time.sleep(1),减轻服务器压力,降低被封IP的风险
  • 异常处理:对requests.get()和soup.find()等操作添加try-except块,避免单条数据爬取失败导致整个程序中断

内容的提问来源于stack exchange,提问作者AK0901

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 21:40:58