You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取多页网站仅获第一页数据求助

问题分析与解决方案

核心问题

你的代码只返回最后一页数据,主要原因是**data = []被放在了页码循环的内部**,每次迭代都会清空列表,最终只保留最后一页爬取到的内容。除此之外还有几个细节问题导致数据异常:

  • 价格字段映射错误:把旧价格赋值给了"New Prices",新价格赋值给了"Old Prices",完全搞反了
  • 未添加请求头,目标网站可能拦截无标识的爬虫请求,导致返回非预期页面
  • 使用zip会在四个元素列表长度不一致时自动截断,丢失部分产品数据
  • 未处理请求失败的异常情况

修正后的代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 初始化数据列表,必须放在循环外面
data = []
name_selector = ".name"
old_price_selector = ".old"
new_price_selector = ".prc"

# 添加请求头,模拟浏览器访问,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 爬取1到49页(range(1,50)对应1-49,共49页)
for page_num in range(1, 50):
    url = f"https://www.jumia.com.ng/phones-tablets/samsung/?q=samsung+phones&page={page_num}#catalog-listing"
    try:
        response = requests.get(url, headers=headers, timeout=10)
        # 检查请求是否成功
        response.raise_for_status()
        soup = BeautifulSoup(response.content, 'html.parser')
        
        # 先定位所有产品卡片容器,再逐个提取信息,避免元素长度不匹配
        products = soup.select(".prd._fb.col.c-prd")
        
        for product in products:
            # 逐个提取单产品信息,缺失字段用"N/A"填充
            name = product.select_one(name_selector).get_text(strip=True) if product.select_one(name_selector) else "N/A"
            old_price = product.select_one(old_price_selector).get_text(strip=True) if product.select_one(old_price_selector) else "N/A"
            new_price = product.select_one(new_price_selector).get_text(strip=True) if product.select_one(new_price_selector) else "N/A"
            discount = product.find("div", {"class": "bdg _dsct _sm"}).get_text(strip=True) if product.find("div", {"class": "bdg _dsct _sm"}) else "N/A"
            
            data.append({
                "Phone Names": name,
                "Old Prices": old_price,
                "New Prices": new_price,
                "Discounts": discount
            })
        print(f"第{page_num}页爬取完成,当前累计{len(data)}条数据")
    except Exception as e:
        print(f"第{page_num}页爬取失败:{str(e)}")
        continue

# 生成DataFrame
df = pd.DataFrame(data)
# 可选:保存为CSV文件
# df.to_csv("samsung_phones_jumia.csv", index=False, encoding="utf-8-sig")

关键改进点

  • 将data = []移到循环外部,确保所有页面的数据都被追加到同一个列表中
  • 先定位产品容器再提取信息,避免抓取到页面无关元素,同时解决列表长度不一致的问题
  • 添加User-Agent请求头,模拟浏览器访问,降低被反爬拦截的概率
  • 加入异常处理,遇到请求失败时跳过当前页,继续爬取后续页面
  • 修正价格字段的映射错误
  • 使用get_text(strip=True)去除文本多余空格和换行,让数据更整洁

内容的提问来源于stack exchange,提问作者Chioma Amuwa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 04:24:32