Python循环爬取books.toscrape.com仅得20条数据,求问题排查与解决
爬虫问题排查:爬取books.toscrape.com仅获取20条数据
我是Python网页爬虫新手,正在尝试爬取books.toscrape.com网站的全部50页数据(需提取内容包括页码、价格、评分、书名、图片URL),并导出为CSV文件。我写了循环遍历页面的代码,但最终CSV只有20行数据,远低于预期的约1000条。怀疑问题出在页面URL循环部分,但不确定具体原因,附上完整代码恳请帮忙排查解决。
# 需要提取的数据: # - 页码 # - 价格 # - 评分 # - 书名 # - 图片URL import bs4 from bs4 import BeautifulSoup import requests import pandas as pd import requests # 创建空列表用于存储提取的数据 pagesList=[] pricesList=[] ratingsList=[] titleList=[] urlsList=[] # 存储爬取数据的字典 book_data={'Title':titleList,'Price':pricesList,'Ratings':ratingsList,'URL':urlsList} # 要爬取的总页数 no_of_pages = 50 # 循环生成所有页面的URL for i in range(1,no_of_pages+1): # 包含最后一页 url=('https://books.toscrape.com/catalogue/page-{}.html'.format(i)) pagesList.append(url) # 将页面URL加入列表 print("页面数量: ",len(pagesList)) print(pagesList) # 请求页面并转为BeautifulSoup对象 for item in pagesList: page=requests.get(item) soup=bs4.BeautifulSoup(page.text,'html.parser') # 格式化输出soup内容 print(soup.prettify()) # 提取书名并加入列表 for t in soup.findAll('h3'): titles=t.getText() titleList.append(titles) # 提取价格并加入列表 for p in soup.find_all('p', class_='price_color'): price=p.getText() pricesList.append(price) # 提取评分并加入列表 for s in soup.find_all('p', class_='star-rating'): for k,v in s.attrs.items(): # k是class,v是star-rating相关值 star=v[1] # 获取评分的字符串值 ratingsList.append(star) print(star) # 提取图片URL并加入列表 divs=soup.find_all('div', class_='image_container') for thumbs in divs: tags=thumbs.find('img', class_='thumbnail') links='https://books.toscrape.com/' + str(tags['src']) newlinks=links.replace('..','') # 去除多余的.. urlsList.append(newlinks) # 存储数据的字典 web_data={'Title':titleList,'Price':pricesList,'Ratings':ratingsList,'URL':urlsList} # 检查各列表长度是否一致(否则Pandas无法正常生成DataFrame) print(len(titleList)) print(len(pricesList)) print(len(ratingsList)) print(len(urlsList)) # 转换为Pandas DataFrame df=pd.DataFrame(web_data) # 修改索引从1开始 df.index+=1 # 移除价格中的货币符号 df['Price']=df['Price'].str.replace('£','') # 按价格降序排序 df.sort_values(by='Price',ascending=False, inplace = True) # 将评分字符串转为对应整数 df['Ratings']=df['Ratings'].replace({'Three':3,'One':1,'Two':2,'Four':4,'Five':5}) # 检查数据类型 df.dtypes # 将价格列转为float类型 df['Price']=df['Price'].astype(float) df.dtypes # 导出为CSV文件 df.to_csv('bookstore.csv')
问题根源
核心错误是数据提取代码没有包含在页面遍历循环内:你只在循环里请求了所有页面并更新soup对象,但循环结束后soup只保留最后一页的内容,后续的书名、价格等提取操作都只针对最后一页执行,所以只得到20条数据(该网站每页正好20本书)。
修复步骤
- 将所有数据提取代码(书名、价格、评分、图片URL的循环)缩进,放入页面遍历的
for item in pagesList:循环内部,确保每请求一页就提取该页的所有数据。 - 可以去掉单独的
pagesList,直接在循环中生成URL并处理,简化代码。 - 新增页码数据的提取(你需求里提到要页码,但原代码没实现)。
修改后的完整代码
# 需要提取的数据: # - 页码 # - 价格 # - 评分 # - 书名 # - 图片URL import bs4 from bs4 import BeautifulSoup import requests import pandas as pd # 创建空列表用于存储提取的数据 pagesList=[] pricesList=[] ratingsList=[] titleList=[] urlsList=[] # 要爬取的总页数 no_of_pages = 50 # 循环遍历每一页 for page_num in range(1, no_of_pages + 1): # 生成当前页的URL url = f'https://books.toscrape.com/catalogue/page-{page_num}.html' # 请求页面 response = requests.get(url) response.raise_for_status() # 捕获请求错误 # 转为BeautifulSoup对象 soup = bs4.BeautifulSoup(response.text, 'html.parser') # 提取当前页的所有书籍数据 books = soup.find_all('article', class_='product_pod') for book in books: # 页码 pagesList.append(page_num) # 书名 title = book.h3.a['title'] titleList.append(title) # 价格 price = book.find('p', class_='price_color').get_text() pricesList.append(price) # 评分 rating_classes = book.find('p', class_='star-rating')['class'] rating = rating_classes[1] # 提取评分等级字符串 ratingsList.append(rating) # 图片URL img_src = book.find('img', class_='thumbnail')['src'] img_url = f'https://books.toscrape.com/catalogue/{img_src.replace("../", "")}' urlsList.append(img_url) # 存储数据的字典 web_data = { '页码': pagesList, '书名': titleList, '价格': pricesList, '评分': ratingsList, '图片URL': urlsList } # 检查各列表长度是否一致 print(f"页码列表长度: {len(pagesList)}") print(f"书名列表长度: {len(titleList)}") print(f"价格列表长度: {len(pricesList)}") print(f"评分列表长度: {len(ratingsList)}") print(f"图片URL列表长度: {len(urlsList)}") # 转换为Pandas DataFrame df = pd.DataFrame(web_data) # 移除价格中的货币符号并转为float df['价格'] = df['价格'].str.replace('£', '').astype(float) # 将评分字符串转为对应整数 rating_mapping = {'One':1, 'Two':2, 'Three':3, 'Four':4, 'Five':5} df['评分'] = df['评分'].map(rating_mapping) # 按页码排序(可选,保留爬取顺序) df.sort_values(by='页码', inplace=True) # 修改索引从1开始 df.index += 1 # 导出为CSV文件 df.to_csv('bookstore_full.csv', index=False, encoding='utf-8-sig') print("数据导出完成,共爬取{}条数据".format(len(df)))
修复说明
- 将数据提取逻辑放入页面循环内,确保每一页的内容都被处理。
- 直接通过
product_pod类定位每本书的容器,让数据提取更精准。 - 新增了页码数据的提取,满足你的需求。
- 添加
response.raise_for_status()捕获请求异常,避免因页面请求失败导致数据缺失。 - 优化图片URL的拼接方式,避免多余的字符替换错误。
- 导出CSV时指定
encoding='utf-8-sig',避免中文乱码问题。
内容的提问来源于stack exchange,提问作者Jtaylor44t
相关产品推荐
相关产品推荐

