Python 3 TypeError排查:爬取Daraz数据时range超20报错
问题描述
爬取Daraz.com.bd手机壳数据时,当循环范围设为range(1,61)会触发类型错误,即使在Google Colab中也无法运行;但将循环改为range(1,21)或range(21,41)这类小范围区间时,代码能正常执行。
错误信息
TypeError: the JSON object must be str, bytes or bytearray, not Tag
完整错误栈
Traceback (most recent call last): File "/Users/fz/Documents/test/scraping_scrpit.py", line 25, in <module> items = json.loads(script)['mods']['listItems'] File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/json/__init__.py", line 339, in loads raise TypeError(f'the JSON object must be str, bytes or bytearray, ' TypeError: the JSON object must be str, bytes or bytearray, not Tag
爬取代码
import requests import json from bs4 import BeautifulSoup as bs import pandas as pd import matplotlib.pyplot as plt import numpy as np headers = { 'User-Agent' : 'Mozilla/5.0' } main_url = 'https://www.daraz.com.bd/' search_url = 'mobile-cases-covers' category_links = [] results = [] for x in range(1,61): url = main_url+search_url+'/?page='+str(x) res = requests.get(url, headers = headers) soup = bs(res.content, 'lxml') # print(res) for script in soup.select('script'): if 'window.pageData=' in script.text: script = script.text.replace('window.pageData=','') break items = json.loads(script)['mods']['listItems'] print(x) for item in items: #print(item) #extract other info you want row = [item['name'], item['inStock'], item['priceShow'], item['price'], item['productUrl'], item['ratingScore'], item['review'], item['cheapest_sku'], item['description'], item['brandId'], item['brandName'], item['sellerName']] results.append(row) df = pd.DataFrame(results, columns = ['Name', 'instock', 'Price show', 'price', 'ProductUrl', 'Rating', 'review', 'cheapest_sku', 'description', 'brandId','brandName','sellerName']) df.to_csv(r"/Users/fz/Documents/test/mobile-cases-covers_1.csv", encoding='utf-8', index=False)
问题原因与解决方案
原因分析
- 反爬机制触发:短时间内连续请求60页数据,网站会识别为爬虫行为,返回的页面结构发生变化,找不到包含
window.pageData=的script标签。此时script变量仍是BeautifulSoup的Tag对象,而非处理后的JSON字符串,传给json.loads()就会触发类型错误。 - 缺失容错逻辑:代码默认每次都能找到目标script标签,但如果没找到,循环结束后
script仍是遍历到的最后一个Tag对象,直接调用json.loads()必然报错。
解决方案
添加异常捕获
在JSON解析环节加入try-except块,捕获类型错误和键错误,跳过出错页面并记录信息,避免程序崩溃:try: items = json.loads(script)['mods']['listItems'] print(x) # 后续数据提取逻辑 except (TypeError, KeyError) as e: print(f"第{x}页爬取失败: {str(e)}") continue降低请求频率
在每次请求后添加延迟,减少被反爬识别的概率:import time # ... res = requests.get(url, headers=headers) time.sleep(1) # 间隔1秒再处理下一页优化目标script查找
明确判断是否找到目标脚本,没找到直接跳过当前页面:target_script = None for script in soup.select('script'): if 'window.pageData=' in script.text: target_script = script.text.replace('window.pageData=', '') break if not target_script: print(f"第{x}页未找到目标数据") continue items = json.loads(target_script)['mods']['listItems']检查请求状态
解析页面前先判断请求是否成功,避免处理错误页面:res = requests.get(url, headers=headers) if res.status_code != 200: print(f"第{x}页请求失败,状态码: {res.status_code}") continue
内容的提问来源于stack exchange,提问作者Fahim Zaman Anik
相关产品推荐
相关产品推荐

