You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从HTML中提取JavaScript函数参数并整理为字典

实现步骤与代码示例

1. 依赖准备

先安装解析HTML所需的库:

pip install beautifulsoup4

2. 核心实现逻辑

  • 用BeautifulSoup解析HTML,筛选出所有href属性包含javascript:fnStep2的<a>标签
  • 用正则表达式提取fnStep2(...)内的参数内容,清理格式后整理为列表
  • 初始化覆盖01到31所有日期的字典,默认值设为空列表
  • 从参数的日期字符串中截取最后两位作为键,将参数列表填充到对应日期的键值下

3. 完整代码

from bs4 import BeautifulSoup
import re

def extract_fnstep2_params(html_content):
    # 初始化包含01-31日期的字典,默认空列表
    date_params = {f"{d:02d}": [] for d in range(1, 32)}
    
    # 解析HTML文档
    soup = BeautifulSoup(html_content, "html.parser")
    # 筛选目标a标签
    target_links = soup.find_all("a", href=re.compile(r"javascript:fnStep2\("))
    
    # 匹配函数参数的正则
    param_pattern = re.compile(r"javascript:fnStep2\((.*?)\)")
    
    for link in target_links:
        href = link.get("href")
        match = param_pattern.search(href)
        if match:
            params_str = match.group(1)
            # 分割参数并清理引号和空格
            params = [p.strip().strip("'\"") for p in params_str.split(",")]
            # 提取日期键(取第一个参数的最后两位)
            if params:
                date_key = params[0][-2:]
                if date_key in date_params:
                    date_params[date_key] = params
    
    return date_params

# 示例使用
sample_html = '''
<a href="javascript:fnStep2('20221001','210','7','0','Y')">1号</a>
<a href="javascript:fnStep2('20221002','210','1','0','Y')">2号</a>
<a href="javascript:fnStep3('20221003')">3号</a>
'''
result = extract_fnstep2_params(sample_html)
print(result)

4. 代码说明

  • 正则表达式采用非贪婪匹配,避免同时匹配多个函数调用的内容
  • 参数处理时自动去除前后引号和空格,保证参数格式整洁
  • 初始化字典直接覆盖1-31号,确保无对应参数的日期也能保留空列表的默认值

内容的提问来源于stack exchange,提问作者shim robin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 01:20:31