使用BeautifulSoup爬取多页网站仅获第一页数据求助
问题分析与解决方案
核心问题
你的代码只返回最后一页数据,主要原因是**data = []被放在了页码循环的内部**,每次迭代都会清空列表,最终只保留最后一页爬取到的内容。除此之外还有几个细节问题导致数据异常:
- 价格字段映射错误:把旧价格赋值给了"New Prices",新价格赋值给了"Old Prices",完全搞反了
- 未添加请求头,目标网站可能拦截无标识的爬虫请求,导致返回非预期页面
- 使用
zip会在四个元素列表长度不一致时自动截断,丢失部分产品数据 - 未处理请求失败的异常情况
修正后的代码
import requests from bs4 import BeautifulSoup import pandas as pd # 初始化数据列表,必须放在循环外面 data = [] name_selector = ".name" old_price_selector = ".old" new_price_selector = ".prc" # 添加请求头,模拟浏览器访问,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 爬取1到49页(range(1,50)对应1-49,共49页) for page_num in range(1, 50): url = f"https://www.jumia.com.ng/phones-tablets/samsung/?q=samsung+phones&page={page_num}#catalog-listing" try: response = requests.get(url, headers=headers, timeout=10) # 检查请求是否成功 response.raise_for_status() soup = BeautifulSoup(response.content, 'html.parser') # 先定位所有产品卡片容器,再逐个提取信息,避免元素长度不匹配 products = soup.select(".prd._fb.col.c-prd") for product in products: # 逐个提取单产品信息,缺失字段用"N/A"填充 name = product.select_one(name_selector).get_text(strip=True) if product.select_one(name_selector) else "N/A" old_price = product.select_one(old_price_selector).get_text(strip=True) if product.select_one(old_price_selector) else "N/A" new_price = product.select_one(new_price_selector).get_text(strip=True) if product.select_one(new_price_selector) else "N/A" discount = product.find("div", {"class": "bdg _dsct _sm"}).get_text(strip=True) if product.find("div", {"class": "bdg _dsct _sm"}) else "N/A" data.append({ "Phone Names": name, "Old Prices": old_price, "New Prices": new_price, "Discounts": discount }) print(f"第{page_num}页爬取完成,当前累计{len(data)}条数据") except Exception as e: print(f"第{page_num}页爬取失败:{str(e)}") continue # 生成DataFrame df = pd.DataFrame(data) # 可选:保存为CSV文件 # df.to_csv("samsung_phones_jumia.csv", index=False, encoding="utf-8-sig")
关键改进点
- 将
data = []移到循环外部,确保所有页面的数据都被追加到同一个列表中 - 先定位产品容器再提取信息,避免抓取到页面无关元素,同时解决列表长度不一致的问题
- 添加
User-Agent请求头,模拟浏览器访问,降低被反爬拦截的概率 - 加入异常处理,遇到请求失败时跳过当前页,继续爬取后续页面
- 修正价格字段的映射错误
- 使用
get_text(strip=True)去除文本多余空格和换行,让数据更整洁
内容的提问来源于stack exchange,提问作者Chioma Amuwa
相关产品推荐
相关产品推荐

