You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页抓取:如何用Try/Except处理缺失值

Fixing "arrays must all be same length" Error When Building Restaurant Data DataFrame

Hey there, let's work through this length mismatch issue you're hitting. The core problem is your websites list is 5 elements shorter than the other data lists (71 vs 76), which breaks pandas' requirement for equal-length arrays when creating a DataFrame. Here are two solid solutions to fix this:

The best approach is to ensure every list gets an entry every time you process a restaurant—even if the website is missing. This way, all lists stay in sync from the start, avoiding mismatched lengths entirely.

Use a try/except block (or a simple existence check) when extracting the website, and append a placeholder like None if the data is missing. Here's how that might look in your scraping code:

# Initialize all your data lists
restaurant_names = []
restaurant_addresses = []
restaurant_websites = []

# Loop through each restaurant entry you're scraping
for entry in restaurant_scrape_results:
    # Extract required fields (assuming these are always present)
    restaurant_names.append(entry.find("h2", class_="name").text.strip())
    restaurant_addresses.append(entry.find("div", class_="address").text.strip())
    
    # Handle the optional website field with try/except
    try:
        website_link = entry.find("a", class_="website-link").get("href")
        restaurant_websites.append(website_link)
    except AttributeError:
        # If the website element doesn't exist, append None
        restaurant_websites.append(None)

This ensures every iteration adds one element to all lists, so their lengths stay identical throughout the scrape.

2. Retroactively Pad the Shorter List

If you've already finished scraping and just need to fix the existing lists, you can pad the shorter websites list with None (or another placeholder like "N/A") until it matches the length of the other lists.

First, confirm the required length (use the length of one of your longer lists), then calculate how many placeholders you need:

# Get the target length from one of the full lists
target_length = len(restaurant_names)

# Calculate how many missing entries we need to add
missing_entries = target_length - len(restaurant_websites)

# Pad the websites list with None
restaurant_websites += [None] * missing_entries

After this, all lists will have the same length, and you can safely create your DataFrame with pd.DataFrame({"name": restaurant_names, "address": restaurant_addresses, "website": restaurant_websites}).

Key Notes to Remember

  • Avoid data misalignment: The retro padding method works only if the order of entries in websites matches exactly with the other lists. If your scrape skipped some restaurants entirely (instead of just missing their websites), this could mislink data—so the first approach is always safer.
  • Placeholder choice: Using None lets pandas automatically convert those entries to NaN, which is easy to work with for missing value analysis later. If you prefer a human-readable placeholder, you can use "No Website Found" instead.
  • Debug first: Always print out list lengths before building the DataFrame to spot mismatches early:
    print(f"Names length: {len(restaurant_names)}")
    print(f"Addresses length: {len(restaurant_addresses)}")
    print(f"Websites length: {len(restaurant_websites)}")
    

内容的提问来源于stack exchange,提问作者lfo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:57:37