You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python循环爬取books.toscrape.com仅得20条数据,求问题排查与解决

爬虫问题排查:爬取books.toscrape.com仅获取20条数据

我是Python网页爬虫新手,正在尝试爬取books.toscrape.com网站的全部50页数据(需提取内容包括页码、价格、评分、书名、图片URL),并导出为CSV文件。我写了循环遍历页面的代码,但最终CSV只有20行数据,远低于预期的约1000条。怀疑问题出在页面URL循环部分,但不确定具体原因,附上完整代码恳请帮忙排查解决。

# 需要提取的数据:
# - 页码
# - 价格
# - 评分
# - 书名
# - 图片URL

import bs4
from bs4 import BeautifulSoup
import requests
import pandas as pd
import requests


# 创建空列表用于存储提取的数据
pagesList=[]
pricesList=[]
ratingsList=[]
titleList=[]
urlsList=[]

# 存储爬取数据的字典
book_data={'Title':titleList,'Price':pricesList,'Ratings':ratingsList,'URL':urlsList}

# 要爬取的总页数
no_of_pages = 50

# 循环生成所有页面的URL
for i in range(1,no_of_pages+1): # 包含最后一页
    url=('https://books.toscrape.com/catalogue/page-{}.html'.format(i))
    pagesList.append(url) # 将页面URL加入列表

print("页面数量: ",len(pagesList))
print(pagesList)

# 请求页面并转为BeautifulSoup对象
for item in pagesList:
    page=requests.get(item)
    soup=bs4.BeautifulSoup(page.text,'html.parser') 
    
# 格式化输出soup内容
print(soup.prettify())

# 提取书名并加入列表
for t in soup.findAll('h3'):
    titles=t.getText()
    titleList.append(titles)

# 提取价格并加入列表
for p in soup.find_all('p', class_='price_color'):
    price=p.getText()
    pricesList.append(price)
    
# 提取评分并加入列表
for s in soup.find_all('p', class_='star-rating'):
    for k,v in s.attrs.items(): # k是class,v是star-rating相关值
        star=v[1] # 获取评分的字符串值
        ratingsList.append(star)
        print(star)

# 提取图片URL并加入列表
divs=soup.find_all('div', class_='image_container')
for thumbs in divs:
    tags=thumbs.find('img', class_='thumbnail')
    links='https://books.toscrape.com/' + str(tags['src'])
    newlinks=links.replace('..','') # 去除多余的..
    urlsList.append(newlinks)

# 存储数据的字典
web_data={'Title':titleList,'Price':pricesList,'Ratings':ratingsList,'URL':urlsList}

# 检查各列表长度是否一致(否则Pandas无法正常生成DataFrame)
print(len(titleList))
print(len(pricesList))
print(len(ratingsList))
print(len(urlsList))

# 转换为Pandas DataFrame
df=pd.DataFrame(web_data)

# 修改索引从1开始
df.index+=1

# 移除价格中的货币符号
df['Price']=df['Price'].str.replace('£','')

# 按价格降序排序
df.sort_values(by='Price',ascending=False, inplace = True)

# 将评分字符串转为对应整数
df['Ratings']=df['Ratings'].replace({'Three':3,'One':1,'Two':2,'Four':4,'Five':5})

# 检查数据类型
df.dtypes

# 将价格列转为float类型
df['Price']=df['Price'].astype(float)
df.dtypes

# 导出为CSV文件
df.to_csv('bookstore.csv')

问题根源

核心错误是数据提取代码没有包含在页面遍历循环内:你只在循环里请求了所有页面并更新soup对象,但循环结束后soup只保留最后一页的内容,后续的书名、价格等提取操作都只针对最后一页执行,所以只得到20条数据(该网站每页正好20本书)。

修复步骤

  1. 将所有数据提取代码(书名、价格、评分、图片URL的循环)缩进,放入页面遍历的for item in pagesList:循环内部,确保每请求一页就提取该页的所有数据。
  2. 可以去掉单独的pagesList,直接在循环中生成URL并处理,简化代码。
  3. 新增页码数据的提取(你需求里提到要页码,但原代码没实现)。

修改后的完整代码

# 需要提取的数据:
# - 页码
# - 价格
# - 评分
# - 书名
# - 图片URL

import bs4
from bs4 import BeautifulSoup
import requests
import pandas as pd

# 创建空列表用于存储提取的数据
pagesList=[]
pricesList=[]
ratingsList=[]
titleList=[]
urlsList=[]

# 要爬取的总页数
no_of_pages = 50

# 循环遍历每一页
for page_num in range(1, no_of_pages + 1):
    # 生成当前页的URL
    url = f'https://books.toscrape.com/catalogue/page-{page_num}.html'
    # 请求页面
    response = requests.get(url)
    response.raise_for_status()  # 捕获请求错误
    # 转为BeautifulSoup对象
    soup = bs4.BeautifulSoup(response.text, 'html.parser')
    
    # 提取当前页的所有书籍数据
    books = soup.find_all('article', class_='product_pod')
    for book in books:
        # 页码
        pagesList.append(page_num)
        # 书名
        title = book.h3.a['title']
        titleList.append(title)
        # 价格
        price = book.find('p', class_='price_color').get_text()
        pricesList.append(price)
        # 评分
        rating_classes = book.find('p', class_='star-rating')['class']
        rating = rating_classes[1]  # 提取评分等级字符串
        ratingsList.append(rating)
        # 图片URL
        img_src = book.find('img', class_='thumbnail')['src']
        img_url = f'https://books.toscrape.com/catalogue/{img_src.replace("../", "")}'
        urlsList.append(img_url)

# 存储数据的字典
web_data = {
    '页码': pagesList,
    '书名': titleList,
    '价格': pricesList,
    '评分': ratingsList,
    '图片URL': urlsList
}

# 检查各列表长度是否一致
print(f"页码列表长度: {len(pagesList)}")
print(f"书名列表长度: {len(titleList)}")
print(f"价格列表长度: {len(pricesList)}")
print(f"评分列表长度: {len(ratingsList)}")
print(f"图片URL列表长度: {len(urlsList)}")

# 转换为Pandas DataFrame
df = pd.DataFrame(web_data)

# 移除价格中的货币符号并转为float
df['价格'] = df['价格'].str.replace('£', '').astype(float)

# 将评分字符串转为对应整数
rating_mapping = {'One':1, 'Two':2, 'Three':3, 'Four':4, 'Five':5}
df['评分'] = df['评分'].map(rating_mapping)

# 按页码排序(可选,保留爬取顺序)
df.sort_values(by='页码', inplace=True)

# 修改索引从1开始
df.index += 1

# 导出为CSV文件
df.to_csv('bookstore_full.csv', index=False, encoding='utf-8-sig')

print("数据导出完成,共爬取{}条数据".format(len(df)))

修复说明

  • 将数据提取逻辑放入页面循环内,确保每一页的内容都被处理。
  • 直接通过product_pod类定位每本书的容器,让数据提取更精准。
  • 新增了页码数据的提取,满足你的需求。
  • 添加response.raise_for_status()捕获请求异常,避免因页面请求失败导致数据缺失。
  • 优化图片URL的拼接方式,避免多余的字符替换错误。
  • 导出CSV时指定encoding='utf-8-sig',避免中文乱码问题。

内容的提问来源于stack exchange,提问作者Jtaylor44t

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 08:25:01