You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

亚马逊客户评论爬取:缺失数据致DataFrame转换失败求助

Hey there! Let's work through this frustrating row mismatch error you're getting when converting scraped Amazon reviews to a DataFrame. This is super common with real-world web data, especially when some fields (like author info, where you see "Ein Kunde" for anonymous users) are missing for certain reviews.

Core Fix Principle

The root cause is that your scraped field lists (e.g., authors, ratings, review texts) have different lengths—some reviews have all fields, others skip missing ones. To fix this, you need to ensure every field list has exactly the same number of elements, using placeholder values (like None or a custom marker) for missing data.


Method 1: Initialize Equal-Length Lists & Force Field Population

Start by creating empty lists for every field you want to scrape. For each review, regardless of whether the field exists, add a value to the corresponding list (use a placeholder if missing). This guarantees all lists stay the same length.

Here's a concrete example tailored to your "Ein Kunde" scenario:

import pandas as pd
# 假设你用BeautifulSoup爬取,先初始化空列表
authors = []
ratings = []
review_texts = []

# 遍历页面上的每条评论元素
for review in soup.find_all("div", class_="review"):
    # 处理作者字段:把"Ein Kunde"当作缺失/匿名处理
    author_elem = review.find("span", class_="a-profile-name")
    if author_elem and author_elem.text.strip() != "Ein Kunde":
        authors.append(author_elem.text.strip())
    else:
        authors.append("Anonymous")  # 或者用None,根据你的需求选择
    
    # 处理星级字段
    rating_elem = review.find("i", class_="review-rating")
    ratings.append(rating_elem.text.split()[0] if rating_elem else None)
    
    # 处理评论内容
    review_elem = review.find("span", class_="review-text-content")
    review_texts.append(review_elem.text.strip() if review_elem else None)

# 现在转换DataFrame不会报错了
df = pd.DataFrame({
    "Author": authors,
    "Rating": ratings,
    "Review Text": review_texts
})

Method 2: Store Individual Reviews as Dictionaries

This approach is more readable: create a dictionary for each review, populating all fields (using placeholders for missing ones), then collect all dictionaries into a list and convert it to a DataFrame. Pandas automatically handles missing values in dictionaries gracefully.

Example:

import pandas as pd
review_data = []

for review in soup.find_all("div", class_="review"):
    review_dict = {}
    
    # 作者字段处理
    author_elem = review.find("span", class_="a-profile-name")
    review_dict["Author"] = (
        author_elem.text.strip() 
        if (author_elem and author_elem.text.strip() != "Ein Kunde") 
        else "Anonymous"
    )
    
    # 星级字段处理
    rating_elem = review.find("i", class_="review-rating")
    review_dict["Rating"] = rating_elem.text.split()[0] if rating_elem else None
    
    # 评论内容处理
    review_elem = review.find("span", class_="review-text-content")
    review_dict["Review Text"] = review_elem.text.strip() if review_elem else None
    
    review_data.append(review_dict)

# 直接转换字典列表为DataFrame
df = pd.DataFrame(review_data)

Method 3: Post-Scrape List Padding (Not Recommended)

If you already have uneven-length lists and don't want to re-scrape, you can pad shorter lists with placeholders to match the longest list's length. Note: This is risky because you might misalign data (e.g., a missing author gets paired with the wrong review). Use only as a last resort.

max_length = max(len(authors), len(ratings), len(review_texts))

# 补全短列表
authors += ["Anonymous"] * (max_length - len(authors))
ratings += [None] * (max_length - len(ratings))
review_texts += [None] * (max_length - len(review_texts))

df = pd.DataFrame({
    "Author": authors,
    "Rating": ratings,
    "Review Text": review_texts
})

Extra Tips

  • Always test your selectors on reviews with missing fields to make sure they correctly return None instead of failing silently.
  • For "Ein Kunde", you can choose to keep it as-is instead of replacing it with "Anonymous"—adjust the logic based on whether you want to distinguish anonymous users from truly missing author data.

内容的提问来源于stack exchange,提问作者lisah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:37:07