如何从HTML的script标签中提取__moduleData__变量爬取Lazada商品评论
提取script标签内__moduleData__变量的解决方法
你可以通过字符串切割+JSON解析的方式获取目标数据,不需要引入额外的JS解析库,具体实现如下:
步骤说明
- 遍历所有script标签定位包含
var __moduleData__的目标标签,避免用固定索引[123]导致页面结构变动后代码失效 - 对script标签内的文本做切割,提取变量赋值语句中等于号后的JSON结构,删除末尾分号得到合法JSON字符串
- 用Python内置的
json库将字符串解析为字典,按层级取到评论数组后写入CSV即可
完整可运行代码
import requests from bs4 import BeautifulSoup import csv import json # 用with语句打开文件自动管理句柄,指定编码避免特殊字符乱码 with open("ProductReviews.csv", "a", encoding="utf-8", newline="") as a_file: writer = csv.writer(a_file) # 写入CSV表头 writer.writerow(["created_at", "reviewer_name", "rating", "content", "source"]) url = 'https://www.lazada.com.my/products/iron-gym-total-upper-body-workout-bar-i467342383.html' # 加UA请求头避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 请求URL response = requests.get(url, headers=headers) # 解析HTML存入BeautifulSoup对象 soup = BeautifulSoup(response.content, "html.parser") # 遍历查找目标script标签 target_script = None for script in soup.find_all("script"): if script.string and 'var __moduleData__' in script.string: target_script = script.string break if target_script: # 切割出JSON字符串:去掉前缀变量定义,去掉尾部的分号 json_str = target_script.split('var __moduleData__ =')[-1].strip().rstrip(';') # 解析为Python字典 module_data = json.loads(json_str) # 按层级提取评论列表 reviews = module_data['data']['root']['fields']['review']['reviews'] # 遍历写入CSV for review in reviews: writer.writerow([ review['reviewTime'], review['reviewer'], review['rating'], review['reviewContent'], 'Lazada' ])
内容的提问来源于stack exchange,提问作者Abdullah Ajmal
相关产品推荐
相关产品推荐

