Python网页抓取:如何用Try/Except处理缺失值
Hey there, let's work through this length mismatch issue you're hitting. The core problem is your websites list is 5 elements shorter than the other data lists (71 vs 76), which breaks pandas' requirement for equal-length arrays when creating a DataFrame. Here are two solid solutions to fix this:
1. Fix the Issue at the Source (Recommended)
The best approach is to ensure every list gets an entry every time you process a restaurant—even if the website is missing. This way, all lists stay in sync from the start, avoiding mismatched lengths entirely.
Use a try/except block (or a simple existence check) when extracting the website, and append a placeholder like None if the data is missing. Here's how that might look in your scraping code:
# Initialize all your data lists restaurant_names = [] restaurant_addresses = [] restaurant_websites = [] # Loop through each restaurant entry you're scraping for entry in restaurant_scrape_results: # Extract required fields (assuming these are always present) restaurant_names.append(entry.find("h2", class_="name").text.strip()) restaurant_addresses.append(entry.find("div", class_="address").text.strip()) # Handle the optional website field with try/except try: website_link = entry.find("a", class_="website-link").get("href") restaurant_websites.append(website_link) except AttributeError: # If the website element doesn't exist, append None restaurant_websites.append(None)
This ensures every iteration adds one element to all lists, so their lengths stay identical throughout the scrape.
2. Retroactively Pad the Shorter List
If you've already finished scraping and just need to fix the existing lists, you can pad the shorter websites list with None (or another placeholder like "N/A") until it matches the length of the other lists.
First, confirm the required length (use the length of one of your longer lists), then calculate how many placeholders you need:
# Get the target length from one of the full lists target_length = len(restaurant_names) # Calculate how many missing entries we need to add missing_entries = target_length - len(restaurant_websites) # Pad the websites list with None restaurant_websites += [None] * missing_entries
After this, all lists will have the same length, and you can safely create your DataFrame with pd.DataFrame({"name": restaurant_names, "address": restaurant_addresses, "website": restaurant_websites}).
Key Notes to Remember
- Avoid data misalignment: The retro padding method works only if the order of entries in
websitesmatches exactly with the other lists. If your scrape skipped some restaurants entirely (instead of just missing their websites), this could mislink data—so the first approach is always safer. - Placeholder choice: Using
Nonelets pandas automatically convert those entries toNaN, which is easy to work with for missing value analysis later. If you prefer a human-readable placeholder, you can use"No Website Found"instead. - Debug first: Always print out list lengths before building the DataFrame to spot mismatches early:
print(f"Names length: {len(restaurant_names)}") print(f"Addresses length: {len(restaurant_addresses)}") print(f"Websites length: {len(restaurant_websites)}")
内容的提问来源于stack exchange,提问作者lfo

