如何用Selenium从页面Script标签提取指定车辆信息?
解决Script标签内车辆信息提取问题
核心思路
跳过框架切换的误区,直接定位包含目标数据的Script标签,提取并解析其中的JS对象数据。以下是具体实现步骤:
1. 确保页面渲染完成
点击查询按钮后,必须等待数据加载完毕(避免过早查找导致元素未出现):
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 点击查询按钮后,等待Script标签加载完成 WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.TAG_NAME, "script")) )
2. 定位目标Script标签
无需纠结Controller框架,先切回主文档(如果之前在AppWindow框架),遍历所有Script标签筛选含目标数据的内容:
# 切回主文档上下文 driver.switch_to.default_content() # 遍历所有Script标签,找到包含车辆数据的那个 target_content = None for script in driver.find_elements(By.TAG_NAME, "script"): inner_html = script.get_attribute("innerHTML") # 用目标字段作为筛选标识 if all(key in inner_html for key in ["Marque et type", "Dénomination commerciale"]): target_content = inner_html break
3. 解析Script中的数据
多数情况下,数据是以JS对象形式存储的,通过字符串截取+JSON解析提取字段:
import json if target_content: # 截取JSON格式的内容(从第一个{到最后一个}) start_idx = target_content.find("{") end_idx = target_content.rfind("}") + 1 json_raw = target_content[start_idx:end_idx] # 处理可能的单引号转双引号(适配JSON解析要求) json_raw = json_raw.replace("'", "\"") # 解析为字典 vehicle_info = json.loads(json_raw) # 提取所需字段 marque_type = vehicle_info.get("Marque et type") denomination = vehicle_info.get("Dénomination commerciale") variante = vehicle_info.get("Variante") version = vehicle_info.get("Version") # 输出结果 print(f"Marque et type: {marque_type}") print(f"Dénomination commerciale: {denomination}") print(f"Variante: {variante}") print(f"Version: {version}")
4. 针对Controller框架的备选方案
如果确认数据确实在Controller框架内,调整框架切换的时机(必须等待框架加载完成):
# 点击查询后等待Controller框架可用并切换 WebDriverWait(driver, 10).until( EC.frame_to_be_available_and_switch_to_it((By.NAME, "Controller")) ) # 后续Script查找、解析步骤与上述一致
内容的提问来源于stack exchange,提问作者Abdelmoula
相关产品推荐
相关产品推荐

