You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+BeautifulSoup遍历多页:如何跳过不存在元素避免报错?

Fixing AttributeError When Missing License Elements in Your Web Scraper

Hey there! I see exactly what's going wrong here. When your script hits a page without that Badge-Condition paragraph (or the <a> tag inside it), trying to call .text or access the href attribute directly throws an AttributeError—you're trying to access properties on a None value, which Python doesn't allow.

Let's adjust your script to handle these missing elements gracefully. We'll check if each element exists before pulling data from it, and set empty values when they're missing. Here's the revised code:

from urllib.request import urlopen
from bs4 import BeautifulSoup as soup
import json

base_url = "https://open.umn.edu/opentextbooks/"
data = []
n = 50

for i in range(4, n+1):
    try:
        response = urlopen(base_url + "BookDetail.aspx?bookId=" + str(i))
        page_html = response.read()
        response.close()
    except Exception as e:
        print(f"Failed to fetch page {i}: {e}")
        continue  # Skip this page if we can't load it
    
    # HTML parsing
    page_soup = soup(page_html, "html.parser")
    containers = page_soup.findAll("div", {"class": "LongDescription"})
    author = page_soup.select("p")
    
    for container in containers:
        item = {}
        item['type'] = "Textbook"
        
        # Safely get title
        title_container = container.find("div", {"class": "twothird"})
        item['title'] = title_container.h1.text.strip() if title_container and title_container.h1 else "No Title"
        
        # Handle author with fallback
        if len(author) >= 4:
            author_text = author[3].get_text(separator=', ').strip()
            item['author'] = author_text if author_text else "University of Minnesota Libraries Publishing"
        else:
            item['author'] = "University of Minnesota Libraries Publishing"
        
        item['link'] = f"{base_url}BookDetail.aspx?bookId={i}"
        item['source'] = "Open Textbook Library"
        item['base_url'] = base_url
        
        # Safely get license info (the main fix!)
        badge_condition = container.find("p", {"class": "Badge-Condition"})
        if badge_condition and badge_condition.a:
            item['license'] = badge_condition.a.text.strip()
            item['license_url'] = badge_condition.a.get("href", "")
        else:
            item['license'] = ""
            item['license_url'] = ""
        
        data.append(item)

# Save cleaned data to JSON
with open("./json/noSubject/otl-loop.json", "w") as writeJSON:
    json.dump(data, writeJSON, ensure_ascii=False, indent=2)

Key Fixes & Improvements:

  • Safe element checks: For every piece of data we extract (title, license, etc.), we first verify the element exists before accessing its properties. No more trying to call .text on None!
  • Page fetch error handling: Wrapped the urlopen call in a try/except block so if a page fails to load entirely, the script skips it instead of crashing.
  • Cleaner author logic: Checks if the author text is actually meaningful before using the fallback value.
  • Whitespace cleanup: Used strip() to remove extra spaces from text values, keeping your data tidy.
  • Simplified URL formatting: Switched to an f-string for building the book link—it's easier to read and maintain.

Now when your script hits a page without license information, it'll just leave those fields empty and keep running through the rest of the pages.

内容的提问来源于stack exchange,提问作者brawlins4

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:46:01