You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用Pandas提取含JavaScript的HTML文件目标值失败

解决方法

问题根源

pandas.read_html()仅能解析静态HTML内容,你需要的「Tempo Total: 03:09:09」是通过JavaScript动态计算渲染生成的,静态解析时只会读取到JS变量定义(如verificar tempo),而非最终渲染后的结果。

可行解决方案

1. 浏览器渲染引擎(推荐批量处理)

用playwright或selenium模拟浏览器加载页面,等待JS执行完成后提取内容,适配复杂渲染逻辑:

  • 安装依赖:
    pip install playwright pandas
    playwright install chromium
    
  • 示例代码:
    import pandas as pd
    from playwright.sync_api import sync_playwright
    import os
    
    def get_tempo(file_path):
        with sync_playwright() as p:
            browser = p.chromium.launch(headless=True)
            page = browser.new_page()
            page.goto(f"file://{os.path.abspath(file_path)}")
            # 替换为实际显示Tempo Total的元素选择器(比如类名、ID)
            tempo_el = page.wait_for_selector(".tempo-total")
            tempo_text = tempo_el.text_content().strip()
            browser.close()
            return tempo_text
    
    # 批量处理HTML文件
    html_dir = "你的HTML文件所在目录"
    files = [f for f in os.listdir(html_dir) if f.endswith(".html")]
    dataset = []
    for fname in files:
        full_path = os.path.join(html_dir, fname)
        tempo = get_tempo(full_path)
        dataset.append({"filename": fname, "tempo_total": tempo})
    
    df = pd.DataFrame(dataset)
    print(df)
    

2. 直接解析JS变量(适用于简单逻辑)

如果verificar tempo变量直接存储了目标时间值,可通过正则表达式提取,无需渲染:

  • 示例代码:
    import pandas as pd
    import re
    import os
    
    def extract_tempo_from_js(file_path):
        with open(file_path, "r", encoding="utf-8") as f:
            content = f.read()
        # 匹配变量赋值语句,根据实际JS语法调整正则
        pattern = r'verificar\s+tempo\s*=\s*"(Tempo Total: \d{2}:\d{2}:\d{2})"'
        match = re.search(pattern, content)
        return match.group(1) if match else None
    
    # 批量处理逻辑同上
    html_dir = "你的HTML文件所在目录"
    files = [f for f in os.listdir(html_dir) if f.endswith(".html")]
    dataset = []
    for fname in files:
        full_path = os.path.join(html_dir, fname)
        tempo = extract_tempo_from_js(full_path)
        dataset.append({"filename": fname, "tempo_total": tempo})
    
    df = pd.DataFrame(dataset)
    

3. 轻量JS渲染(requests-html)

requests-html内置简单渲染引擎,适合快速测试:

  • 安装依赖:
    pip install requests-html pandas
    
  • 示例代码:
    import pandas as pd
    from requests_html import HTMLSession
    import os
    
    def get_tempo(file_path):
        session = HTMLSession()
        res = session.get(f"file://{os.path.abspath(file_path)}")
        res.html.render()  # 触发JS渲染
        tempo_text = res.html.find(".tempo-total", first=True).text.strip()
        session.close()
        return tempo_text
    
    # 批量处理逻辑一致
    

注意点

  • 浏览器渲染方案批量处理时,建议控制并发数(比如用线程池),避免资源占用过高;
  • 正则解析需根据实际JS代码调整匹配规则,确保精准度;
  • 优先选择浏览器渲染方案,兼容性更强,能应对复杂JS逻辑。

内容的提问来源于stack exchange,提问作者Dini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 19:23:23