Python+BeautifulSoup遍历多页:如何跳过不存在元素避免报错?
Hey there! I see exactly what's going wrong here. When your script hits a page without that Badge-Condition paragraph (or the <a> tag inside it), trying to call .text or access the href attribute directly throws an AttributeError—you're trying to access properties on a None value, which Python doesn't allow.
Let's adjust your script to handle these missing elements gracefully. We'll check if each element exists before pulling data from it, and set empty values when they're missing. Here's the revised code:
from urllib.request import urlopen from bs4 import BeautifulSoup as soup import json base_url = "https://open.umn.edu/opentextbooks/" data = [] n = 50 for i in range(4, n+1): try: response = urlopen(base_url + "BookDetail.aspx?bookId=" + str(i)) page_html = response.read() response.close() except Exception as e: print(f"Failed to fetch page {i}: {e}") continue # Skip this page if we can't load it # HTML parsing page_soup = soup(page_html, "html.parser") containers = page_soup.findAll("div", {"class": "LongDescription"}) author = page_soup.select("p") for container in containers: item = {} item['type'] = "Textbook" # Safely get title title_container = container.find("div", {"class": "twothird"}) item['title'] = title_container.h1.text.strip() if title_container and title_container.h1 else "No Title" # Handle author with fallback if len(author) >= 4: author_text = author[3].get_text(separator=', ').strip() item['author'] = author_text if author_text else "University of Minnesota Libraries Publishing" else: item['author'] = "University of Minnesota Libraries Publishing" item['link'] = f"{base_url}BookDetail.aspx?bookId={i}" item['source'] = "Open Textbook Library" item['base_url'] = base_url # Safely get license info (the main fix!) badge_condition = container.find("p", {"class": "Badge-Condition"}) if badge_condition and badge_condition.a: item['license'] = badge_condition.a.text.strip() item['license_url'] = badge_condition.a.get("href", "") else: item['license'] = "" item['license_url'] = "" data.append(item) # Save cleaned data to JSON with open("./json/noSubject/otl-loop.json", "w") as writeJSON: json.dump(data, writeJSON, ensure_ascii=False, indent=2)
Key Fixes & Improvements:
- Safe element checks: For every piece of data we extract (title, license, etc.), we first verify the element exists before accessing its properties. No more trying to call
.textonNone! - Page fetch error handling: Wrapped the
urlopencall in a try/except block so if a page fails to load entirely, the script skips it instead of crashing. - Cleaner author logic: Checks if the author text is actually meaningful before using the fallback value.
- Whitespace cleanup: Used
strip()to remove extra spaces from text values, keeping your data tidy. - Simplified URL formatting: Switched to an f-string for building the book link—it's easier to read and maintain.
Now when your script hits a page without license information, it'll just leave those fields empty and keep running through the rest of the pages.
内容的提问来源于stack exchange,提问作者brawlins4

