You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python BeautifulSoup抓取Prestashop网站的混合JS/JSON内容?

解决Prestashop网站混合JS/JSON数据的爬取与解析问题

核心思路

从包含JavaScript的script标签中提取纯JSON结构,清理掉不符合JSON规范的语法后解析,再将数据存入数据库。

具体步骤与代码实现

1. 提取目标Script标签内容

先通过BeautifulSoup定位到包含产品数据的script标签(通常在页面底部):

import requests
from bs4 import BeautifulSoup

url = "目标批发商网站的产品页URL"
# 添加请求头模拟浏览器,避免被拦截
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

# 找到包含目标数据的script标签(根据实际内容特征调整筛选条件)
script_tag = soup.find('script', text=lambda text: text and ('prestashop' in text or 'product' in text))
if not script_tag:
    print("未找到包含目标数据的script标签")
    exit()
script_content = script_tag.string

2. 提取类JSON字符串

用正则匹配出JS变量赋值后的JSON结构(根据实际响应调整变量名):

import re

# 匹配var prestashop = 后面的对象内容,直到分号结束(re.DOTALL允许匹配多行)
match_result = re.search(r'var prestashop = (.*?);', script_content, re.DOTALL)
if not match_result:
    # 如果变量名不是prestashop,换成实际的变量名重试
    match_result = re.search(r'var product = (.*?);', script_content, re.DOTALL)
    if not match_result:
        print("未匹配到目标JSON结构")
        exit()
raw_json_str = match_result.group(1)

3. 清理非JSON语法

JS中的函数、undefined、末尾逗号等不符合JSON规范,需要先清理:

# 移除所有函数定义(比如function(){...})
cleaned_json_str = re.sub(r'function\s*\([^)]*\)\s*{[^}]*}', '', raw_json_str)
# 将undefined替换为JSON支持的null
cleaned_json_str = cleaned_json_str.replace('undefined', 'null')
# 移除对象/数组末尾的逗号(比如{key: value,} → {key: value})
cleaned_json_str = re.sub(r',\s*([}\]])', r'\1', cleaned_json_str)

4. 解析JSON数据

用Python内置的json模块解析清理后的字符串:

import json

try:
    product_data = json.loads(cleaned_json_str)
    # 打印解析后的核心数据,验证是否正确
    print("解析成功,产品ID:", product_data['product']['id'])
    print("产品名称:", product_data['product']['name'])
except json.JSONDecodeError as e:
    print(f"JSON解析失败:{e}")
    # 若解析失败,打印清理后的字符串排查问题
    # print(cleaned_json_str)

5. 存入数据库(以SQLite为例)

用轻量级的SQLite存储数据,适合新手快速上手:

import sqlite3

# 连接数据库(不存在则自动创建)
conn = sqlite3.connect('prestashop_products.db')
cursor = conn.cursor()

# 创建产品表(根据实际字段调整,比如添加价格、库存等)
cursor.execute('''
CREATE TABLE IF NOT EXISTS products (
    id INTEGER PRIMARY KEY,
    name TEXT NOT NULL,
    price REAL,
    description TEXT,
    stock_quantity INTEGER
)
''')

# 提取需要存入的字段(根据解析后的data结构调整)
target_product = product_data['product']
product_info = (
    target_product['id'],
    target_product['name'],
    target_product['price'],
    target_product['description'],
    target_product['quantity']
)

# 插入数据
cursor.execute('''
INSERT INTO products (id, name, price, description, stock_quantity)
VALUES (?, ?, ?, ?, ?)
''', product_info)

# 提交更改并关闭连接
conn.commit()
conn.close()
print("数据已成功存入数据库")

新手注意事项

  • 调整正则匹配规则:不同Prestashop网站的变量名可能不同,打印script_content查看实际变量名后修改正则。
  • 排查解析错误:如果出现JSONDecodeError,打印cleaned_json_str查看剩余的不符合JSON的内容,针对性添加清理规则。
  • 批量爬取注意反爬:如果爬取多个页面,添加请求间隔(time.sleep(2)),避免触发网站反爬机制。

内容的提问来源于stack exchange,提问作者yasar avci

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 01:09:59