You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium加载本地HTML过慢,Playwright依赖问题求助

本地解析Pro-Football-Reference HTML文件的性能与依赖问题

我从pro-football-reference.com爬取数据时因访问限制被封禁,所以把所有HTML页面下载到本地data文件夹。现在遇到三个核心问题:

一、Selenium解析本地文件速度极慢

用Selenium打开500KB左右的本地HTML文件,耗时约95秒,代码如下:

def Parse_File(box_score_file):
    options = Options()
    options.headless = True
    driver = webdriver.Chrome(options=options)
    driver.get(f'file://{box_score_file}')
    html = driver.page_source
        
    soup = BeautifulSoup(html, 'html.parser')
    [s.decompose() for s in soup.select('tr.over_header')]
    [s.decompose() for s in soup.select('tr.thead')]
    return soup

二、Requests无法获取JS渲染的内容

尝试用requests直接读取本地文件,但部分表格由JavaScript动态生成/隐藏,无法提取完整数据生成DataFrame。

三、Playwright依赖报错

尝试改用Playwright,但始终报错ModuleNotFoundError: No module named 'pyee.asyncio'。已执行pip install pyee,Anaconda显示依赖已满足,同步/异步方式均出现该错误,代码如下:

from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeout

def Parse_File(box_score_file):
    try:
        with sync_playwright() as p:
            browser = p.chromium.launch()
            page = browser.new_page()
            page.goto(box_score_file)
            print(page.title())
            html = page.inner_html()
    except PlaywrightTimeout:
        print(f"Timeout error on {box_score_file}")
        
    soup = BeautifulSoup(html, 'html.parser')
    [s.decompose() for s in soup.select('tr.over_header')]
    [s.decompose() for s in soup.select('tr.thead')]
    return soup

优化与修复方案

针对Selenium性能问题

  1. 复用浏览器实例:每次调用函数都启动新Chrome实例是主要耗时点,将浏览器初始化移到函数外复用:
# 全局初始化一次浏览器
options = Options()
options.headless = True
driver = webdriver.Chrome(options=options)

def Parse_File(box_score_file):
    driver.get(f'file://{box_score_file}')
    html = driver.page_source
        
    soup = BeautifulSoup(html, 'html.parser')
    [s.decompose() for s in soup.select('tr.over_header')]
    [s.decompose() for s in soup.select('tr.thead')]
    return soup

# 所有文件解析完成后关闭浏览器
# driver.quit()
  1. 禁用非必要资源加载:关闭图片加载、限制JS执行,减少页面加载耗时:
options.add_argument('--blink-settings=imagesEnabled=false')
# 仅保留必要JS,若目标表格不需要JS渲染可设为2禁用
options.add_experimental_option("prefs", {
    "profile.managed_default_content_settings.javascript": 1
})

针对Playwright依赖问题

  1. 完全重装Playwright:卸载现有包后用官方命令重装,避免环境冲突:
pip uninstall -y playwright pyee
pip install playwright
playwright install chromium
  1. 使用纯Python虚拟环境:避开Anaconda的依赖冲突,创建独立环境测试:
python -m venv pfr_env
source pfr_env/bin/activate  # Linux/macOS
# pfr_env\Scripts\activate  # Windows
pip install playwright beautifulsoup4
playwright install chromium

备选方案:直接解析本地HTML并修复隐藏内容

如果仅部分表格被JS隐藏(如带display:none或hidden类),可直接用BeautifulSoup读取文件后手动解锁隐藏元素,无需启动浏览器:

def Parse_File(box_score_file):
    with open(box_score_file, 'r', encoding='utf-8') as f:
        html = f.read()
    
    soup = BeautifulSoup(html, 'html.parser')
    # 移除隐藏属性,显示被JS隐藏的内容
    for elem in soup.select('[style*="display:none"], .hidden'):
        elem.attrs.pop('style', None)
        elem.attrs.pop('class', None)
    
    [s.decompose() for s in soup.select('tr.over_header')]
    [s.decompose() for s in soup.select('tr.thead')]
    return soup

该方法速度最快,需根据实际HTML结构调整隐藏元素的选择器。


内容的提问来源于stack exchange,提问作者Rain Maker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 23:55:18