You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取遇AttributeError:'NoneType'无find_all属性,求多页爬取方案

网页爬取问题求助

我是网页爬取新手,使用.find_all时遇到问题。能成功爬取https://books.toscrape.com/首页,但爬取第二页或批量爬取全部50页时,都会抛出AttributeError: 'NoneType' object has no attribute 'find_all'错误,错误出在提取书籍分类的代码行,相关代码及报错信息如下,求解决批量爬取50页的问题:

爬取第二页的代码

response = requests.get(url)
soup = BeautifulSoup(response.content,'html.parser')

#find all the book titles and their links under h3 tag
books = soup.find_all('h3')

book_extracted = 0

#iterate through the books and extract the information of each book
for book in books:
    book_url = book.find('a')['href'] #grabbing or acessing href attribute of 1st book
    book_response = requests.get(url + book_url) #getting the response of the 1st book link
    book_soup = BeautifulSoup(book_response.content,'html.parser')
    
    
    title = book_soup.find('h1').text #extracting title of the 1st book
    category = book_soup.find('ul',class_='breadcrumb').find_all('a')[2].text.strip() 
    rating = book_soup.find('p',class_='star-rating')['class'][1] 
    price = book_soup.find('p',class_='price_color').text.strip()
    availability = book_soup.find('p',class_='availability').text.strip()
    
    book_extracted += 1
    
    print(f"Title: {title}")
    print(f"Category:{category}")
    print(f"Rating:{rating}")
    print(f"Price:{price}")
    print(f"Availability:{availability}")
    print("***************")

爬取第二页的报错信息

AttributeError                            Traceback (most recent call last)
Cell In[80], line 18
     14 book_soup = BeautifulSoup(book_response.content,'html.parser')
     17 title = book_soup.find('h1').text #extracting title of the 1st book
---> 18 category = book_soup.find('ul',class_='breadcrumb').find_all('a')[2].text.strip() #extracting category
     19 rating = book_soup.find('p',class_='star-rating')['class'][1] #having two classes star_rating and rating(three)
     20 price = book_soup.find('p',class_='price_color').text.strip()

AttributeError: 'NoneType' object has no attribute 'find_all'

批量爬取50页的代码

books_data = []

#loop through all 50 pages
for page_num in range(1,51):
    url = f'https://books.toscrape.com/catalogue/page-{page_num}.html'
    response = requests.get(url)
    soup = BeautifulSoup(response.content,'html.parser')
    
    # find all the book titles and their links under h3 tag of the current page
    books = soup.find_all('h3')
    
    for book in books:
        book_url = book.find('a')['href'] #grabbing or acessing href attribute of 1st book
        book_response = requests.get(url + book_url) #getting the response of the 1st book link
        book_soup = BeautifulSoup(book_response.content,'html.parser')
    
        title = book_soup.find('h1').text #extracting title of the 1st book
        category = book_soup.find('ul',class_='breadcrumb').find_all('a')[2].text.strip() 
        rating = book_soup.find('p',class_='star-rating')['class'][1] 
        price = book_soup.find('p',class_='price_color').text.strip()
        availability = book_soup.find('p',class_='availability').text.strip()
    
        #appending the extracted data to the list
        books_data.append([title,category,rating,price,availability])
        print(books_data)

批量爬取50页的报错信息

AttributeError                            Traceback (most recent call last)
Cell In[82], line 20
     16 book_soup = BeautifulSoup(book_response.content,'html.parser')
     19 title = book_soup.find('h1').text #extracting title of the 1st book
---> 20 category = book_soup.find('ul',class_='breadcrumb').find_all('a')[2].text.strip() #extracting category
     21 rating = book_soup.find('p',class_='star-rating')['class'][1] #having two classes star_rating and rating(three)
     22 price = book_soup.find('p',class_='price_color').text.strip()

AttributeError: 'NoneType' object has no attribute 'find_all'

问题原因与解决方法

核心问题:URL拼接错误

你在拼接书籍详情页URL时犯了路径错误:

  • 第二页的URL是https://books.toscrape.com/catalogue/page-2.html,而书籍的href值是类似a-light-in-the-attic_1000/index.html的相对路径。直接用url + book_url会得到无效的URL(比如https://books.toscrape.com/catalogue/page-2.htmla-light-in-the-attic_1000/index.html),请求返回的不是正常的书籍详情页,导致book_soup.find('ul',class_='breadcrumb')找不到目标元素,返回None,调用.find_all就会触发报错。

修复步骤

  1. 修正URL拼接逻辑:
    书籍详情页的基础路径是https://books.toscrape.com/catalogue/,应该用这个基础路径加上书籍的相对href,而不是当前页的完整URL。

  2. 添加异常处理(可选但推荐):
    加入try-except块捕获异常,避免单个书籍爬取失败导致整个程序中断,同时可以记录错误信息。

修复后的批量爬取代码

import requests
from bs4 import BeautifulSoup
import time

books_data = []
base_catalogue_url = 'https://books.toscrape.com/catalogue/'

# 遍历50页
for page_num in range(1, 51):
    page_url = f'{base_catalogue_url}page-{page_num}.html'
    response = requests.get(page_url)
    soup = BeautifulSoup(response.content, 'html.parser')
    
    books = soup.find_all('h3')
    
    for book in books:
        book_relative_url = book.find('a')['href']
        # 正确拼接书籍详情页URL
        book_full_url = base_catalogue_url + book_relative_url
        
        try:
            book_response = requests.get(book_full_url)
            book_response.raise_for_status()  # 检查请求是否成功
            book_soup = BeautifulSoup(book_response.content, 'html.parser')
            
            title = book_soup.find('h1').text.strip()
            # 先判断breadcrumb是否存在,避免None调用方法
            breadcrumb = book_soup.find('ul', class_='breadcrumb')
            category = breadcrumb.find_all('a')[2].text.strip() if breadcrumb else '未知分类'
            rating = book_soup.find('p', class_='star-rating')['class'][1]
            price = book_soup.find('p', class_='price_color').text.strip()
            availability = book_soup.find('p', class_='availability').text.strip()
            
            books_data.append([title, category, rating, price, availability])
            print(f"已爬取:{title}")
            
            # 添加请求延迟,避免触发反爬
            time.sleep(0.5)
            
        except Exception as e:
            print(f"爬取失败:{book_full_url},错误:{str(e)}")

# 可选:将数据保存为CSV文件
import csv
with open('books_data.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.writer(f)
    writer.writerow(['标题', '分类', '评分', '价格', '库存'])
    writer.writerows(books_data)

额外建议

  • 请求延迟:添加time.sleep()降低请求频率,避免被网站限制访问。
  • Session复用:使用requests.Session()建立持久连接,提升爬取效率。
  • 状态码检查:用response.raise_for_status()检查请求是否成功,及时发现无效请求。

内容的提问来源于stack exchange,提问作者Maddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 14:45:08