如何用Python BeautifulSoup抓取Prestashop网站的混合JS/JSON内容?
解决Prestashop网站混合JS/JSON数据的爬取与解析问题
核心思路
从包含JavaScript的script标签中提取纯JSON结构,清理掉不符合JSON规范的语法后解析,再将数据存入数据库。
具体步骤与代码实现
1. 提取目标Script标签内容
先通过BeautifulSoup定位到包含产品数据的script标签(通常在页面底部):
import requests from bs4 import BeautifulSoup url = "目标批发商网站的产品页URL" # 添加请求头模拟浏览器,避免被拦截 headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # 找到包含目标数据的script标签(根据实际内容特征调整筛选条件) script_tag = soup.find('script', text=lambda text: text and ('prestashop' in text or 'product' in text)) if not script_tag: print("未找到包含目标数据的script标签") exit() script_content = script_tag.string
2. 提取类JSON字符串
用正则匹配出JS变量赋值后的JSON结构(根据实际响应调整变量名):
import re # 匹配var prestashop = 后面的对象内容,直到分号结束(re.DOTALL允许匹配多行) match_result = re.search(r'var prestashop = (.*?);', script_content, re.DOTALL) if not match_result: # 如果变量名不是prestashop,换成实际的变量名重试 match_result = re.search(r'var product = (.*?);', script_content, re.DOTALL) if not match_result: print("未匹配到目标JSON结构") exit() raw_json_str = match_result.group(1)
3. 清理非JSON语法
JS中的函数、undefined、末尾逗号等不符合JSON规范,需要先清理:
# 移除所有函数定义(比如function(){...}) cleaned_json_str = re.sub(r'function\s*\([^)]*\)\s*{[^}]*}', '', raw_json_str) # 将undefined替换为JSON支持的null cleaned_json_str = cleaned_json_str.replace('undefined', 'null') # 移除对象/数组末尾的逗号(比如{key: value,} → {key: value}) cleaned_json_str = re.sub(r',\s*([}\]])', r'\1', cleaned_json_str)
4. 解析JSON数据
用Python内置的json模块解析清理后的字符串:
import json try: product_data = json.loads(cleaned_json_str) # 打印解析后的核心数据,验证是否正确 print("解析成功,产品ID:", product_data['product']['id']) print("产品名称:", product_data['product']['name']) except json.JSONDecodeError as e: print(f"JSON解析失败:{e}") # 若解析失败,打印清理后的字符串排查问题 # print(cleaned_json_str)
5. 存入数据库(以SQLite为例)
用轻量级的SQLite存储数据,适合新手快速上手:
import sqlite3 # 连接数据库(不存在则自动创建) conn = sqlite3.connect('prestashop_products.db') cursor = conn.cursor() # 创建产品表(根据实际字段调整,比如添加价格、库存等) cursor.execute(''' CREATE TABLE IF NOT EXISTS products ( id INTEGER PRIMARY KEY, name TEXT NOT NULL, price REAL, description TEXT, stock_quantity INTEGER ) ''') # 提取需要存入的字段(根据解析后的data结构调整) target_product = product_data['product'] product_info = ( target_product['id'], target_product['name'], target_product['price'], target_product['description'], target_product['quantity'] ) # 插入数据 cursor.execute(''' INSERT INTO products (id, name, price, description, stock_quantity) VALUES (?, ?, ?, ?, ?) ''', product_info) # 提交更改并关闭连接 conn.commit() conn.close() print("数据已成功存入数据库")
新手注意事项
- 调整正则匹配规则:不同Prestashop网站的变量名可能不同,打印
script_content查看实际变量名后修改正则。 - 排查解析错误:如果出现
JSONDecodeError,打印cleaned_json_str查看剩余的不符合JSON的内容,针对性添加清理规则。 - 批量爬取注意反爬:如果爬取多个页面,添加请求间隔(
time.sleep(2)),避免触发网站反爬机制。
内容的提问来源于stack exchange,提问作者yasar avci
相关产品推荐
相关产品推荐

