如何用Scrapy爬取<script nonce>标签中的餐厅经纬度数据
Scrapy提取PF Chang墨西哥站点餐厅经纬度
目标站点的经纬度数据嵌在页面第一个<script>标签内,以JS变量形式存储。以下是具体实现代码:
import scrapy import re import json class PFChangSpider(scrapy.Spider): name = 'pf_chang_locations' start_urls = ['https://pfchangsmexico.com.mx/ubicaciones/'] def parse(self, response): # 定位目标script标签并提取文本内容 script_content = response.xpath('/html/body/script[1]/text()').get() if not script_content: self.logger.error("无法获取目标script标签内容") return # 匹配包含餐厅地理数据的JS数组(需根据实际源码调整变量名) data_match = re.search(r'var\s+locations\s*=\s*(\[.*?\]);', script_content, re.DOTALL) if not data_match: self.logger.error("未匹配到餐厅数据数组") return # 解析JSON数据 try: locations = json.loads(data_match.group(1)) except json.JSONDecodeError as e: self.logger.error(f"JSON解析失败: {str(e)}") return # 遍历提取每个餐厅的经纬度及其他信息 for loc in locations: yield { '餐厅名称': loc.get('name'), '纬度': loc.get('lat'), '经度': loc.get('lng'), # 可根据需求添加地址、电话等其他字段 }
关键说明
- 正则匹配调整:若页面script内的变量名不是
locations,需查看源码修改正则中的变量名。 - JSON兼容处理:如果JS数组存在单引号、末尾逗号等非标准JSON语法,需先做替换处理,例如:
cleaned_data = data_match.group(1).replace("'", "\"").replace(",]", "]") locations = json.loads(cleaned_data) - 字段验证:实际返回的JSON结构可能有差异,需根据源码确认
lat/lng等字段的准确名称。
内容的提问来源于stack exchange,提问作者prathamesh wadile
相关产品推荐
相关产品推荐

