You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取网页JS代码中images子段数据?已写部分Python代码求帮助

提取JavaScript中countrydataprovider对象的images数据

实现步骤

  • 提取目标script标签内的完整文本内容
  • 通过正则表达式匹配出countrydataprovider对象的完整结构,处理JavaScript与JSON的语法差异
  • 解析成可操作的Python数据结构,提取images字段

完整代码

import requests
from bs4 import BeautifulSoup
import re
import json

url = "https://mausam.imd.gov.in/imd_latest/contents/stationwise-nowcast-warning.php"
html = requests.get(url).content
soup = BeautifulSoup(html, 'html.parser')
script_tag = soup.find('script', attrs={'type': 'text/javascript'})

# 获取script标签内的文本内容
script_content = script_tag.string

# 正则匹配countrydataprovider对象的完整定义
pattern = r'var countrydataprovider = ({.*?});'
match = re.search(pattern, script_content, re.DOTALL)

if match:
    # 提取匹配到的对象字符串,处理JSON不兼容的语法
    json_str = match.group(1)
    # 移除数组/对象末尾多余的逗号(JavaScript允许但JSON不允许)
    json_str = re.sub(r',\s*([}\]])', r'\1', json_str)
    # 解析为Python字典
    country_data = json.loads(json_str)
    # 提取images字段数据
    images_data = country_data.get('images', [])
    
    # 示例:遍历打印数据
    for item in images_data:
        print(f"站点名称: {item['title']}")
        print(f"经纬度: {item['latitude']}, {item['longitude']}")
        print(f"预警信息: {item['description']}\n")
else:
    print("未找到countrydataprovider对象")

关键说明

  • re.DOTALL参数让正则表达式的.能匹配换行符,确保捕获多行的JavaScript对象
  • 末尾逗号处理是必要的,因为JSON语法不允许数组或对象最后一项后加逗号,而JavaScript代码中常出现这种写法
  • 解析完成后,images_data就是包含所有站点标记数据的Python列表,可直接用于后续处理

内容的提问来源于stack exchange,提问作者G.S. J

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 16:45:42