You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何一次性获取全部结果?代码报KeyError: 'One'及输出格式需求

问题分析与解决方案

1. KeyError: 'One' 错误修复

错误根源是星级映射字典的键与页面实际类名大小写不匹配:页面中星级对应的类名是One、Two这类首字母大写格式,但你的map_rating函数里用的是"one"小写键,导致找不到对应值抛出异常。

同时clean_price函数的正则替换逻辑也有问题——用空格替换非数字字符会导致字符串含空格,转换float时会报错,需改为空字符串替换。

修正后的核心函数:

def clean_price(price):
    return float(re.sub("[^0-9.]", "", price))  # 替换为空字符串

def map_rating(rating):
    rating_map = {
        "one": 1,
        "two": 2,
        "three": 3,
        "four": 4,
        "five": 5
    }
    return rating_map[rating.lower()]  # 转为小写后匹配,兼容大小写

2. 一次性获取所有页面结果的实现

网站采用分页展示数据,需循环遍历所有分页链接,直到找不到下一页为止。完整代码如下:

import requests as r
from bs4 import BeautifulSoup
import re

base_url = "https://books.toscrape.com/"
current_url = base_url
all_book_data = []

def clean_price(price):
    return float(re.sub("[^0-9.]", "", price))

def map_rating(rating):
    rating_map = {
        "one": 1,
        "two": 2,
        "three": 3,
        "four": 4,
        "five": 5
    }
    return rating_map[rating.lower()]

def extract_book_data(book_tag):
    title = book_tag.find("h3").find("a")["title"]
    price = book_tag.find("p", attrs={"class": "price_color"}).get_text()
    rating = book_tag.find("p", attrs={"class": "star-rating"})["class"][-1]
    return {
        "title": title,
        "price": clean_price(price),
        "rating": map_rating(rating)
    }

# 循环遍历所有分页
while True:
    resp = r.get(current_url)
    soup = BeautifulSoup(resp.content, "html.parser")
    
    # 提取当前页书籍数据
    book_tags = soup.find_all("article", attrs={"class": "product_pod"})
    page_books = [extract_book_data(tag) for tag in book_tags]
    all_book_data.extend(page_books)
    
    # 检查是否有下一页
    next_link = soup.find("li", attrs={"class": "next"})
    if not next_link:
        break
    next_page_path = next_link.find("a")["href"]
    current_url = base_url + next_page_path

print(f"共爬取到 {len(all_book_data)} 本书籍")
print(all_book_data[:3])  # 打印前3条示例数据

代码说明

  • 分页处理:通过检测页面底部的next按钮,动态拼接下一页URL,直到无下一页时终止循环。
  • 大小写兼容:rating.lower()确保不管页面返回的类名是大写还是小写,都能正确匹配映射字典。
  • 数据收集:用extend方法将每页的书籍数据合并到总列表中,最终得到全量结果。

内容的提问来源于stack exchange,提问作者Anish Thapa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 10:48:26