求助:使用Pandas提取含JavaScript的HTML文件目标值失败
解决方法
问题根源
pandas.read_html()仅能解析静态HTML内容,你需要的「Tempo Total: 03:09:09」是通过JavaScript动态计算渲染生成的,静态解析时只会读取到JS变量定义(如verificar tempo),而非最终渲染后的结果。
可行解决方案
1. 浏览器渲染引擎(推荐批量处理)
用playwright或selenium模拟浏览器加载页面,等待JS执行完成后提取内容,适配复杂渲染逻辑:
- 安装依赖:
pip install playwright pandas playwright install chromium - 示例代码:
import pandas as pd from playwright.sync_api import sync_playwright import os def get_tempo(file_path): with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page() page.goto(f"file://{os.path.abspath(file_path)}") # 替换为实际显示Tempo Total的元素选择器(比如类名、ID) tempo_el = page.wait_for_selector(".tempo-total") tempo_text = tempo_el.text_content().strip() browser.close() return tempo_text # 批量处理HTML文件 html_dir = "你的HTML文件所在目录" files = [f for f in os.listdir(html_dir) if f.endswith(".html")] dataset = [] for fname in files: full_path = os.path.join(html_dir, fname) tempo = get_tempo(full_path) dataset.append({"filename": fname, "tempo_total": tempo}) df = pd.DataFrame(dataset) print(df)
2. 直接解析JS变量(适用于简单逻辑)
如果verificar tempo变量直接存储了目标时间值,可通过正则表达式提取,无需渲染:
- 示例代码:
import pandas as pd import re import os def extract_tempo_from_js(file_path): with open(file_path, "r", encoding="utf-8") as f: content = f.read() # 匹配变量赋值语句,根据实际JS语法调整正则 pattern = r'verificar\s+tempo\s*=\s*"(Tempo Total: \d{2}:\d{2}:\d{2})"' match = re.search(pattern, content) return match.group(1) if match else None # 批量处理逻辑同上 html_dir = "你的HTML文件所在目录" files = [f for f in os.listdir(html_dir) if f.endswith(".html")] dataset = [] for fname in files: full_path = os.path.join(html_dir, fname) tempo = extract_tempo_from_js(full_path) dataset.append({"filename": fname, "tempo_total": tempo}) df = pd.DataFrame(dataset)
3. 轻量JS渲染(requests-html)
requests-html内置简单渲染引擎,适合快速测试:
- 安装依赖:
pip install requests-html pandas - 示例代码:
import pandas as pd from requests_html import HTMLSession import os def get_tempo(file_path): session = HTMLSession() res = session.get(f"file://{os.path.abspath(file_path)}") res.html.render() # 触发JS渲染 tempo_text = res.html.find(".tempo-total", first=True).text.strip() session.close() return tempo_text # 批量处理逻辑一致
注意点
- 浏览器渲染方案批量处理时,建议控制并发数(比如用线程池),避免资源占用过高;
- 正则解析需根据实际JS代码调整匹配规则,确保精准度;
- 优先选择浏览器渲染方案,兼容性更强,能应对复杂JS逻辑。
内容的提问来源于stack exchange,提问作者Dini
相关产品推荐
相关产品推荐

