You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3 TypeError排查:爬取Daraz数据时range超20报错

问题描述

爬取Daraz.com.bd手机壳数据时,当循环范围设为range(1,61)会触发类型错误,即使在Google Colab中也无法运行;但将循环改为range(1,21)或range(21,41)这类小范围区间时,代码能正常执行。

错误信息

TypeError: the JSON object must be str, bytes or bytearray, not Tag

完整错误栈

Traceback (most recent call last):
  File "/Users/fz/Documents/test/scraping_scrpit.py", line 25, in <module>
    items = json.loads(script)['mods']['listItems']
  File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/json/__init__.py", line 339, in loads
    raise TypeError(f'the JSON object must be str, bytes or bytearray, '
TypeError: the JSON object must be str, bytes or bytearray, not Tag

爬取代码

import requests
import json
from bs4 import BeautifulSoup as bs
import pandas as pd
import matplotlib.pyplot as plt
import numpy as np

headers = {
    'User-Agent' : 'Mozilla/5.0'
}

main_url = 'https://www.daraz.com.bd/'
search_url = 'mobile-cases-covers'
category_links = []
results = []
for x in range(1,61):
    url = main_url+search_url+'/?page='+str(x)
    res = requests.get(url, headers = headers)
    soup = bs(res.content, 'lxml')
    # print(res)
    for script in soup.select('script'):
        if 'window.pageData=' in script.text:
            script = script.text.replace('window.pageData=','')
            break
    items = json.loads(script)['mods']['listItems']
    print(x)

    for item in items:
        #print(item)
        #extract other info you want
        row = [item['name'], item['inStock'], item['priceShow'], item['price'], item['productUrl'], item['ratingScore'], item['review'], item['cheapest_sku'], item['description'], item['brandId'], item['brandName'], item['sellerName']]
        results.append(row)       

df = pd.DataFrame(results, columns = ['Name', 'instock', 'Price show', 'price', 'ProductUrl', 'Rating', 'review', 'cheapest_sku', 'description', 'brandId','brandName','sellerName'])


df.to_csv(r"/Users/fz/Documents/test/mobile-cases-covers_1.csv", encoding='utf-8', index=False)
问题原因与解决方案

原因分析

  • 反爬机制触发:短时间内连续请求60页数据,网站会识别为爬虫行为,返回的页面结构发生变化,找不到包含window.pageData=的script标签。此时script变量仍是BeautifulSoup的Tag对象,而非处理后的JSON字符串,传给json.loads()就会触发类型错误。
  • 缺失容错逻辑:代码默认每次都能找到目标script标签,但如果没找到,循环结束后script仍是遍历到的最后一个Tag对象,直接调用json.loads()必然报错。

解决方案

  1. 添加异常捕获
    在JSON解析环节加入try-except块,捕获类型错误和键错误,跳过出错页面并记录信息,避免程序崩溃:

    try:
        items = json.loads(script)['mods']['listItems']
        print(x)
        # 后续数据提取逻辑
    except (TypeError, KeyError) as e:
        print(f"第{x}页爬取失败: {str(e)}")
        continue
    
  2. 降低请求频率
    在每次请求后添加延迟,减少被反爬识别的概率:

    import time
    # ...
    res = requests.get(url, headers=headers)
    time.sleep(1)  # 间隔1秒再处理下一页
    
  3. 优化目标script查找
    明确判断是否找到目标脚本,没找到直接跳过当前页面:

    target_script = None
    for script in soup.select('script'):
        if 'window.pageData=' in script.text:
            target_script = script.text.replace('window.pageData=', '')
            break
    if not target_script:
        print(f"第{x}页未找到目标数据")
        continue
    items = json.loads(target_script)['mods']['listItems']
    
  4. 检查请求状态
    解析页面前先判断请求是否成功,避免处理错误页面:

    res = requests.get(url, headers=headers)
    if res.status_code != 200:
        print(f"第{x}页请求失败,状态码: {res.status_code}")
        continue
    

内容的提问来源于stack exchange,提问作者Fahim Zaman Anik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 22:48:52