使用Selenium解析SSB网站CPI历史表格遇html5lib导入错误求助
解决pd.read_html找不到html5lib及SSB表格最优解析方案
问题描述
我需要解析挪威统计局(SSB)官网页面上的「表1:消费者价格指数,1924年以来的历史指数(2015=100)」表格。已用Selenium编写代码打开该表格,但执行pd.read_html时抛出错误:
ImportError: html5lib not found, please install it
尽管通过pip list确认已安装html5lib(版本1.1),问题仍未解决。附现有代码:
options = Options() url = "https://www.ssb.no/en/priser-og-prisindekser/konsumpriser/statistikk/konsumprisindeksen" driver_no = webdriver.Chrome(options=options, executable_path=mypath) driver_no.get(url) sleep(2) WebDriverWait(driver_no, 20).until(EC.element_to_be_clickable((By.XPATH, '//*[@id="attachment-table-figure-1"]/button'))) elem = driver_no.find_element(By.XPATH, '//*[@id="attachment-table-figure-1"]/button') sleep(2) driver_no.execute_script("arguments[0].scrollIntoView(true);", elem) sleep(2) driver_no.find_element(By.XPATH, '//*[@id="attachment-table-figure-1"]/button').click() df_list = pd.read_html(driver_no.page_source, "html_parser") driver_no.quit()
一、解决html5lib导入错误
1. 修正pd.read_html参数
你错误地将解析器名称作为第二个参数传入,第二个参数是表格匹配字符串,需用flavor参数指定解析器:
df_list = pd.read_html(driver_no.page_source, flavor='html5lib')
2. 重装依赖包
可能存在依赖冲突或安装不完整,执行以下命令彻底重装相关库:
pip uninstall -y html5lib beautifulsoup4 lxml pip install html5lib beautifulsoup4 lxml
3. 验证Python环境
确认运行代码的Python环境与pip list显示的环境一致(比如避免虚拟环境、conda环境的切换冲突)。
二、最优解析表格方案
方案1:直接调用SSB API(推荐,无需Selenium)
SSB提供结构化数据API,直接请求接口可快速获取表格数据,效率远高于Selenium:
import pandas as pd import requests # 对应表1的API接口 api_url = "https://data.ssb.no/api/v0/en/table/08515/" payload = { "query": [], "response": {"format": "json"} } response = requests.post(api_url, json=payload) data = response.json() # 转换为DataFrame columns = [var["label"] for var in data["variables"]] df = pd.DataFrame(data["data"], columns=columns) print(df.head())
方案2:优化Selenium代码
若必须使用Selenium,优化等待逻辑并直接抓取表格元素的HTML,减少干扰:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd options = Options() options.add_argument('--headless=new') # 无头模式提升运行速度 url = "https://www.ssb.no/en/priser-og-prisindekser/konsumpriser/statistikk/konsumprisindeksen" driver_no = webdriver.Chrome(options=options) try: driver_no.get(url) # 等待按钮可点击并执行点击 btn = WebDriverWait(driver_no, 20).until( EC.element_to_be_clickable((By.XPATH, '//*[@id="attachment-table-figure-1"]/button')) ) driver_no.execute_script("arguments[0].click();", btn) # 等待表格加载完成并获取元素 table = WebDriverWait(driver_no, 20).until( EC.presence_of_element_located((By.ID, 'attachment-table-figure-1-table')) ) # 仅用表格自身的HTML解析,避免整个页面的冗余内容干扰 df_list = pd.read_html(table.get_attribute('innerHTML'), flavor='html5lib') df = df_list[0] print(df.head()) finally: driver_no.quit()
内容的提问来源于stack exchange,提问作者asuidncsdk
相关产品推荐
相关产品推荐

