如何一次性获取全部结果?代码报KeyError: 'One'及输出格式需求
问题分析与解决方案
1. KeyError: 'One' 错误修复
错误根源是星级映射字典的键与页面实际类名大小写不匹配:页面中星级对应的类名是One、Two这类首字母大写格式,但你的map_rating函数里用的是"one"小写键,导致找不到对应值抛出异常。
同时clean_price函数的正则替换逻辑也有问题——用空格替换非数字字符会导致字符串含空格,转换float时会报错,需改为空字符串替换。
修正后的核心函数:
def clean_price(price): return float(re.sub("[^0-9.]", "", price)) # 替换为空字符串 def map_rating(rating): rating_map = { "one": 1, "two": 2, "three": 3, "four": 4, "five": 5 } return rating_map[rating.lower()] # 转为小写后匹配,兼容大小写
2. 一次性获取所有页面结果的实现
网站采用分页展示数据,需循环遍历所有分页链接,直到找不到下一页为止。完整代码如下:
import requests as r from bs4 import BeautifulSoup import re base_url = "https://books.toscrape.com/" current_url = base_url all_book_data = [] def clean_price(price): return float(re.sub("[^0-9.]", "", price)) def map_rating(rating): rating_map = { "one": 1, "two": 2, "three": 3, "four": 4, "five": 5 } return rating_map[rating.lower()] def extract_book_data(book_tag): title = book_tag.find("h3").find("a")["title"] price = book_tag.find("p", attrs={"class": "price_color"}).get_text() rating = book_tag.find("p", attrs={"class": "star-rating"})["class"][-1] return { "title": title, "price": clean_price(price), "rating": map_rating(rating) } # 循环遍历所有分页 while True: resp = r.get(current_url) soup = BeautifulSoup(resp.content, "html.parser") # 提取当前页书籍数据 book_tags = soup.find_all("article", attrs={"class": "product_pod"}) page_books = [extract_book_data(tag) for tag in book_tags] all_book_data.extend(page_books) # 检查是否有下一页 next_link = soup.find("li", attrs={"class": "next"}) if not next_link: break next_page_path = next_link.find("a")["href"] current_url = base_url + next_page_path print(f"共爬取到 {len(all_book_data)} 本书籍") print(all_book_data[:3]) # 打印前3条示例数据
代码说明
- 分页处理:通过检测页面底部的
next按钮,动态拼接下一页URL,直到无下一页时终止循环。 - 大小写兼容:
rating.lower()确保不管页面返回的类名是大写还是小写,都能正确匹配映射字典。 - 数据收集:用
extend方法将每页的书籍数据合并到总列表中,最终得到全量结果。
内容的提问来源于stack exchange,提问作者Anish Thapa
相关产品推荐
相关产品推荐

