网页爬取遇AttributeError:'NoneType'无find_all属性,求多页爬取方案
网页爬取问题求助
我是网页爬取新手,使用.find_all时遇到问题。能成功爬取https://books.toscrape.com/首页,但爬取第二页或批量爬取全部50页时,都会抛出AttributeError: 'NoneType' object has no attribute 'find_all'错误,错误出在提取书籍分类的代码行,相关代码及报错信息如下,求解决批量爬取50页的问题:
爬取第二页的代码
response = requests.get(url) soup = BeautifulSoup(response.content,'html.parser') #find all the book titles and their links under h3 tag books = soup.find_all('h3') book_extracted = 0 #iterate through the books and extract the information of each book for book in books: book_url = book.find('a')['href'] #grabbing or acessing href attribute of 1st book book_response = requests.get(url + book_url) #getting the response of the 1st book link book_soup = BeautifulSoup(book_response.content,'html.parser') title = book_soup.find('h1').text #extracting title of the 1st book category = book_soup.find('ul',class_='breadcrumb').find_all('a')[2].text.strip() rating = book_soup.find('p',class_='star-rating')['class'][1] price = book_soup.find('p',class_='price_color').text.strip() availability = book_soup.find('p',class_='availability').text.strip() book_extracted += 1 print(f"Title: {title}") print(f"Category:{category}") print(f"Rating:{rating}") print(f"Price:{price}") print(f"Availability:{availability}") print("***************")
爬取第二页的报错信息
AttributeError Traceback (most recent call last) Cell In[80], line 18 14 book_soup = BeautifulSoup(book_response.content,'html.parser') 17 title = book_soup.find('h1').text #extracting title of the 1st book ---> 18 category = book_soup.find('ul',class_='breadcrumb').find_all('a')[2].text.strip() #extracting category 19 rating = book_soup.find('p',class_='star-rating')['class'][1] #having two classes star_rating and rating(three) 20 price = book_soup.find('p',class_='price_color').text.strip() AttributeError: 'NoneType' object has no attribute 'find_all'
批量爬取50页的代码
books_data = [] #loop through all 50 pages for page_num in range(1,51): url = f'https://books.toscrape.com/catalogue/page-{page_num}.html' response = requests.get(url) soup = BeautifulSoup(response.content,'html.parser') # find all the book titles and their links under h3 tag of the current page books = soup.find_all('h3') for book in books: book_url = book.find('a')['href'] #grabbing or acessing href attribute of 1st book book_response = requests.get(url + book_url) #getting the response of the 1st book link book_soup = BeautifulSoup(book_response.content,'html.parser') title = book_soup.find('h1').text #extracting title of the 1st book category = book_soup.find('ul',class_='breadcrumb').find_all('a')[2].text.strip() rating = book_soup.find('p',class_='star-rating')['class'][1] price = book_soup.find('p',class_='price_color').text.strip() availability = book_soup.find('p',class_='availability').text.strip() #appending the extracted data to the list books_data.append([title,category,rating,price,availability]) print(books_data)
批量爬取50页的报错信息
AttributeError Traceback (most recent call last) Cell In[82], line 20 16 book_soup = BeautifulSoup(book_response.content,'html.parser') 19 title = book_soup.find('h1').text #extracting title of the 1st book ---> 20 category = book_soup.find('ul',class_='breadcrumb').find_all('a')[2].text.strip() #extracting category 21 rating = book_soup.find('p',class_='star-rating')['class'][1] #having two classes star_rating and rating(three) 22 price = book_soup.find('p',class_='price_color').text.strip() AttributeError: 'NoneType' object has no attribute 'find_all'
问题原因与解决方法
核心问题:URL拼接错误
你在拼接书籍详情页URL时犯了路径错误:
- 第二页的URL是
https://books.toscrape.com/catalogue/page-2.html,而书籍的href值是类似a-light-in-the-attic_1000/index.html的相对路径。直接用url + book_url会得到无效的URL(比如https://books.toscrape.com/catalogue/page-2.htmla-light-in-the-attic_1000/index.html),请求返回的不是正常的书籍详情页,导致book_soup.find('ul',class_='breadcrumb')找不到目标元素,返回None,调用.find_all就会触发报错。
修复步骤
修正URL拼接逻辑:
书籍详情页的基础路径是https://books.toscrape.com/catalogue/,应该用这个基础路径加上书籍的相对href,而不是当前页的完整URL。添加异常处理(可选但推荐):
加入try-except块捕获异常,避免单个书籍爬取失败导致整个程序中断,同时可以记录错误信息。
修复后的批量爬取代码
import requests from bs4 import BeautifulSoup import time books_data = [] base_catalogue_url = 'https://books.toscrape.com/catalogue/' # 遍历50页 for page_num in range(1, 51): page_url = f'{base_catalogue_url}page-{page_num}.html' response = requests.get(page_url) soup = BeautifulSoup(response.content, 'html.parser') books = soup.find_all('h3') for book in books: book_relative_url = book.find('a')['href'] # 正确拼接书籍详情页URL book_full_url = base_catalogue_url + book_relative_url try: book_response = requests.get(book_full_url) book_response.raise_for_status() # 检查请求是否成功 book_soup = BeautifulSoup(book_response.content, 'html.parser') title = book_soup.find('h1').text.strip() # 先判断breadcrumb是否存在,避免None调用方法 breadcrumb = book_soup.find('ul', class_='breadcrumb') category = breadcrumb.find_all('a')[2].text.strip() if breadcrumb else '未知分类' rating = book_soup.find('p', class_='star-rating')['class'][1] price = book_soup.find('p', class_='price_color').text.strip() availability = book_soup.find('p', class_='availability').text.strip() books_data.append([title, category, rating, price, availability]) print(f"已爬取:{title}") # 添加请求延迟,避免触发反爬 time.sleep(0.5) except Exception as e: print(f"爬取失败:{book_full_url},错误:{str(e)}") # 可选:将数据保存为CSV文件 import csv with open('books_data.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['标题', '分类', '评分', '价格', '库存']) writer.writerows(books_data)
额外建议
- 请求延迟:添加
time.sleep()降低请求频率,避免被网站限制访问。 - Session复用:使用
requests.Session()建立持久连接,提升爬取效率。 - 状态码检查:用
response.raise_for_status()检查请求是否成功,及时发现无效请求。
内容的提问来源于stack exchange,提问作者Maddy
相关产品推荐
相关产品推荐

