You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取亚马逊商品数据报错All arrays must be of the same length如何解决

报错原因与修复方案

All arrays must be of the same length报错的核心原因是构造DataFrame时,传入的ratings和review两个列表长度不一致,由以下几个代码问题共同导致:

  • 分页URL硬编码错误:请求链接中page=2写死,循环变量没有生效,所有分页请求都会返回第二页内容
  • 列表追加逻辑有漏洞:仅当评分存在时才给ratings列表加元素,仅当评论存在时才给review列表加元素,遇到只有评分没有评论、或只有评论没有评分的商品时,两个列表长度就会出现差值
  • 评论字段选择器不准确:使用通用类名a-size-base定位评论数,会匹配到大量非评论内容,导致提取结果不符合预期

修复后代码

from bs4 import BeautifulSoup
import requests
import pandas as pd

data = []
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:66.0) Gecko/20100101 Firefox/66.0",
    "Accept-Encoding": "gzip, deflate",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "DNT": "1",
    "Connection": "close",
    "Upgrade-Insecure-Requests": "1",
}
for page in range(1, 5):
    # 修正硬编码的page参数
    r = requests.get(
        "https://www.amazon.com/s?k=redmi&page={page}&qid=1631528810&ref=sr_pg_={page}".format(
            page=page
        ),
        headers=headers,
    )
    soup = BeautifulSoup(r.content, "lxml")
    for d in soup.findAll("div", attrs={"class": "s-result-item"}):
        # 统一提取当前商品的两个字段,不存在则赋值为None
        rating = d.find("span", attrs={"class": "a-icon-alt"})
        rating_val = rating.text if rating else None
        # 更换为评论数专属的类名选择器
        review_cnt = d.find("span", class_="a-size-base s-underline-text")
        review_val = review_cnt.text.strip() if review_cnt else None
        # 成对存储两个字段,避免长度不一致
        data.append({"rating": rating_val, "reviews": review_val})

# 直接通过字典列表构造DataFrame,无长度匹配问题
df = pd.DataFrame(data)
df.to_csv("products.csv", index=False, encoding="utf-8")

内容的提问来源于stack exchange,提问作者Arslan Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 12:24:01