You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup多页爬取,仅部分内容写入CSV的问题排查

问题分析与修复方案:无法将多页爬取内容写入CSV

嘿,我帮你检查了代码,发现几个核心问题导致你没法把所有爬取的内容写入CSV,而且你以为的“多页爬取”其实根本没真正生效,咱们一步步拆解问题并修复:

主要问题点

1. 多页爬取逻辑完全没触发

你生成了urls列表来存各页链接,但后续完全没遍历这个列表去爬取每一页,全程只处理了初始的第一页URL。而且URL拼接的格式化语法也错了——'Page-{1}'.format(url,page)这种写法只会生成Page-1、Page-2这样的字符串,不是完整的页面链接,正确的做法应该把基础URL和页码拼接起来。

2. CSV写入的逻辑错误百出

  • 你写的result = title.replace(",","")," + "," + prev_price +"," +current_price +"\n"会生成一个元组,而write()方法需要传入字符串,这不仅会报错,就算不报错也写不出正确内容。
  • f.write(result)放在了产品循环的外面,这意味着哪怕逻辑正确,也只会写入最后一个产品的信息。
  • 表头headers没有加换行符\n,会导致第一行数据和表头连在一起,CSV格式直接乱掉。

3. 其他细节问题

  • 你在循环产品时,用page_soup.findAll获取评分,这会拿到整个页面的所有评分,每次取第一个,导致所有产品的评分都是同一个值,应该用当前的container来查找对应产品的评分。
  • 直接open()文件没指定编码,容易出现特殊字符乱码。
  • 没有处理产品缺少原价、评分的情况,一旦遇到这种产品,代码会直接报错终止。

修复后的完整代码

from bs4 import BeautifulSoup as soup
from urllib.request import urlopen

# 基础URL,注意保留末尾的斜杠,方便拼接页码
base_url = "https://www.newegg.com/Black-Friday-Deals/EventSaleStore/ID-10475/"
filename = "BlackFridayNewegg.csv"

# 使用with语句自动管理文件,不用手动close,更安全
with open(filename, "w", encoding="utf-8") as f:
    # 写入表头并添加换行,避免和第一行数据粘连
    headers = "Product, Previous price, Current price, Rating\n"
    f.write(headers)

    # 遍历1到6页(range(1,7)会生成1、2、3、4、5、6)
    for page_num in range(1, 7):
        # 正确拼接每页的完整URL
        page_url = f"{base_url}Page-{page_num}"
        print(f"正在爬取第 {page_num} 页: {page_url}")
        
        try:
            # 打开页面并解析
            page = urlopen(page_url)
            html = page.read().decode("utf-8")
            page.close()
            page_soup = soup(html, "html.parser")
            # 获取当前页的所有产品容器
            containers = page_soup.findAll("div", {"class": "item-container"})
            
            # 遍历当前页的每个产品
            for container in containers:
                # 提取产品标题,替换逗号避免CSV列错位
                title_container = container.findAll("a", {"class": "item-title"})
                title = title_container[0].text.strip().replace(",", "")
                
                # 提取评分,处理没有评分的情况
                rating_container = container.findAll("span", {"class": "item-rating-num"})
                rating = rating_container[0].text.strip() if rating_container else "无评分"
                
                # 提取原价,处理没有原价的情况
                prev_price_container = container.findAll("li", {"class": "price-was"})
                prev_price = prev_price_container[0].text.strip() if prev_price_container else "无原价"
                
                # 提取当前价格,处理异常情况
                current_price_container = container.findAll("li", {"class": "price-current"})
                current_price = current_price_container[0].text.strip() if current_price_container else "价格未知"
                
                # 打印当前产品信息(可选,用于调试)
                print(f"产品名称: {title}")
                print(f"原价: {prev_price}")
                print(f"现价: {current_price}")
                print(f"评分: {rating}\n")
                
                # 拼接成符合CSV格式的行,写入文件
                csv_line = f"{title},{prev_price},{current_price},{rating}\n"
                f.write(csv_line)
                
        except Exception as e:
            # 捕获异常,避免单个页面出错导致整个程序崩溃
            print(f"爬取第 {page_num} 页时出错: {str(e)}")
            continue

print("所有页面数据已成功写入CSV文件!")

修复关键点说明

  1. 真正实现多页爬取:现在会逐个遍历1-6页,正确拼接每页的完整URL并爬取内容。
  2. CSV写入逻辑修复:
    • 用with语句管理文件,自动处理关闭操作,避免资源泄漏。
    • 在产品循环内部写入每一行数据,确保所有产品都被记录。
    • 用f-string正确拼接CSV行,替换标题中的逗号防止列错位。
  3. 容错与异常处理:对可能缺失的元素(评分、原价)做了判断,添加try-except捕获爬取错误,保证程序稳定运行。
  4. 编码规范:指定encoding="utf-8"打开文件,避免特殊字符乱码。

内容的提问来源于stack exchange,提问作者Yon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 21:42:31