You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup编写网页爬虫时遭遇KeyError: 'href'错误的排查与解决求助

Fixing KeyError: 'href' in Your BS4 Web Crawler

Hey there! Let's work through that KeyError: 'href' you're running into when scaling your crawler to 80 pages. You’re absolutely right about the root cause—your current code assumes every <a> tag in main_data has an href attribute, but that’s not always the case (some links might be empty, or use JavaScript instead of a direct href).

Root Cause

In your transform_subpage function, you check if main_data is non-empty, but you don’t verify that the first <a> tag actually contains an href property. When you hit a page where the selected <a> tag has no href, main_data[0]['href'] throws a KeyError because that key doesn’t exist in the tag’s attributes.

Solution: Make Your transform_subpage Function Robust

Let’s rewrite the function to add a check for the href attribute before trying to access it. We’ll also simplify the logic by setting a default value upfront:

def transform_subpage(soup_subpage):
    main_data = soup_subpage.select("div.entry-content.clearfix > p > a")
    # Set default value first
    subpage_link = {'subpage_link': '--'}
    
    if main_data:
        target_link = main_data[0]
        # Check if the <a> tag actually has an href attribute
        if 'href' in target_link.attrs:
            subpage_link['subpage_link'] = target_link['href']
    
    subpage.append(subpage_link)
    return

What This Does:

  • We start with a default value (--) so we don’t need duplicate code in an else block.
  • We first confirm main_data has elements, then check if the first <a> tag includes the href attribute using 'href' in target_link.attrs.
  • This handles both cases: no <a> tags found, or <a> tags found but without an href.

Bonus: Add Fault Tolerance to Main Page Scraping

While we’re at it, let’s make your main transform function more resilient too. Sometimes pages might have malformed post blocks (missing titles, links, etc.) that could break your crawler. Adding checks for each element will prevent unexpected errors:

def transform(soup):
    for item in soup.find_all('div', class_ = 'post-block-style'):
        # Safely extract title
        title_elem = item.find('h3', {'class': 'post-title'})
        title = title_elem.text.replace('\n','') if title_elem else '--'
        
        # Safely extract link (check both existence of the tag and href attribute)
        link_elem = item.find('a')
        link = link_elem['href'] if (link_elem and 'href' in link_elem.attrs) else '--'
        
        # Safely extract summary
        summary_elem = item.find('p')
        summary = summary_elem.text if summary_elem else '--'
        
        # Safely extract date
        date_elem = item.find('span', {'class': 'post-meta-date'})
        date = date_elem.text.replace('\n','') if date_elem else '--'
        
        news = {
            'title': title,
            'link': link,
            'summary': summary,
            'date': date
        }
        newslist.append(news)
    return

Why This Helps:

  • If a post block is missing any element (like a title or date), the crawler will just use -- instead of crashing.
  • You’ll get a more complete dataset, even if some pages have inconsistent structure.

内容的提问来源于stack exchange,提问作者ppincus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 19:32:28