Selenium加载本地HTML过慢,Playwright依赖问题求助
本地解析Pro-Football-Reference HTML文件的性能与依赖问题
我从pro-football-reference.com爬取数据时因访问限制被封禁,所以把所有HTML页面下载到本地data文件夹。现在遇到三个核心问题:
一、Selenium解析本地文件速度极慢
用Selenium打开500KB左右的本地HTML文件,耗时约95秒,代码如下:
def Parse_File(box_score_file): options = Options() options.headless = True driver = webdriver.Chrome(options=options) driver.get(f'file://{box_score_file}') html = driver.page_source soup = BeautifulSoup(html, 'html.parser') [s.decompose() for s in soup.select('tr.over_header')] [s.decompose() for s in soup.select('tr.thead')] return soup
二、Requests无法获取JS渲染的内容
尝试用requests直接读取本地文件,但部分表格由JavaScript动态生成/隐藏,无法提取完整数据生成DataFrame。
三、Playwright依赖报错
尝试改用Playwright,但始终报错ModuleNotFoundError: No module named 'pyee.asyncio'。已执行pip install pyee,Anaconda显示依赖已满足,同步/异步方式均出现该错误,代码如下:
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeout def Parse_File(box_score_file): try: with sync_playwright() as p: browser = p.chromium.launch() page = browser.new_page() page.goto(box_score_file) print(page.title()) html = page.inner_html() except PlaywrightTimeout: print(f"Timeout error on {box_score_file}") soup = BeautifulSoup(html, 'html.parser') [s.decompose() for s in soup.select('tr.over_header')] [s.decompose() for s in soup.select('tr.thead')] return soup
优化与修复方案
针对Selenium性能问题
- 复用浏览器实例:每次调用函数都启动新Chrome实例是主要耗时点,将浏览器初始化移到函数外复用:
# 全局初始化一次浏览器 options = Options() options.headless = True driver = webdriver.Chrome(options=options) def Parse_File(box_score_file): driver.get(f'file://{box_score_file}') html = driver.page_source soup = BeautifulSoup(html, 'html.parser') [s.decompose() for s in soup.select('tr.over_header')] [s.decompose() for s in soup.select('tr.thead')] return soup # 所有文件解析完成后关闭浏览器 # driver.quit()
- 禁用非必要资源加载:关闭图片加载、限制JS执行,减少页面加载耗时:
options.add_argument('--blink-settings=imagesEnabled=false') # 仅保留必要JS,若目标表格不需要JS渲染可设为2禁用 options.add_experimental_option("prefs", { "profile.managed_default_content_settings.javascript": 1 })
针对Playwright依赖问题
- 完全重装Playwright:卸载现有包后用官方命令重装,避免环境冲突:
pip uninstall -y playwright pyee pip install playwright playwright install chromium
- 使用纯Python虚拟环境:避开Anaconda的依赖冲突,创建独立环境测试:
python -m venv pfr_env source pfr_env/bin/activate # Linux/macOS # pfr_env\Scripts\activate # Windows pip install playwright beautifulsoup4 playwright install chromium
备选方案:直接解析本地HTML并修复隐藏内容
如果仅部分表格被JS隐藏(如带display:none或hidden类),可直接用BeautifulSoup读取文件后手动解锁隐藏元素,无需启动浏览器:
def Parse_File(box_score_file): with open(box_score_file, 'r', encoding='utf-8') as f: html = f.read() soup = BeautifulSoup(html, 'html.parser') # 移除隐藏属性,显示被JS隐藏的内容 for elem in soup.select('[style*="display:none"], .hidden'): elem.attrs.pop('style', None) elem.attrs.pop('class', None) [s.decompose() for s in soup.select('tr.over_header')] [s.decompose() for s in soup.select('tr.thead')] return soup
该方法速度最快,需根据实际HTML结构调整隐藏元素的选择器。
内容的提问来源于stack exchange,提问作者Rain Maker
相关产品推荐
相关产品推荐

