You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium解析SSB网站CPI历史表格遇html5lib导入错误求助

解决pd.read_html找不到html5lib及SSB表格最优解析方案

问题描述

我需要解析挪威统计局(SSB)官网页面上的「表1:消费者价格指数,1924年以来的历史指数(2015=100)」表格。已用Selenium编写代码打开该表格,但执行pd.read_html时抛出错误:

ImportError: html5lib not found, please install it

尽管通过pip list确认已安装html5lib(版本1.1),问题仍未解决。附现有代码:

options = Options()

url = "https://www.ssb.no/en/priser-og-prisindekser/konsumpriser/statistikk/konsumprisindeksen"
driver_no = webdriver.Chrome(options=options, executable_path=mypath)

driver_no.get(url)
sleep(2)
WebDriverWait(driver_no, 20).until(EC.element_to_be_clickable((By.XPATH, '//*[@id="attachment-table-figure-1"]/button')))
elem = driver_no.find_element(By.XPATH, '//*[@id="attachment-table-figure-1"]/button')
sleep(2)
driver_no.execute_script("arguments[0].scrollIntoView(true);", elem)
sleep(2)
driver_no.find_element(By.XPATH, '//*[@id="attachment-table-figure-1"]/button').click()

df_list = pd.read_html(driver_no.page_source, "html_parser")
driver_no.quit()

一、解决html5lib导入错误

1. 修正pd.read_html参数

你错误地将解析器名称作为第二个参数传入,第二个参数是表格匹配字符串,需用flavor参数指定解析器:

df_list = pd.read_html(driver_no.page_source, flavor='html5lib')

2. 重装依赖包

可能存在依赖冲突或安装不完整,执行以下命令彻底重装相关库:

pip uninstall -y html5lib beautifulsoup4 lxml
pip install html5lib beautifulsoup4 lxml

3. 验证Python环境

确认运行代码的Python环境与pip list显示的环境一致(比如避免虚拟环境、conda环境的切换冲突)。


二、最优解析表格方案

方案1:直接调用SSB API(推荐,无需Selenium)

SSB提供结构化数据API,直接请求接口可快速获取表格数据,效率远高于Selenium:

import pandas as pd
import requests

# 对应表1的API接口
api_url = "https://data.ssb.no/api/v0/en/table/08515/"
payload = {
    "query": [],
    "response": {"format": "json"}
}

response = requests.post(api_url, json=payload)
data = response.json()

# 转换为DataFrame
columns = [var["label"] for var in data["variables"]]
df = pd.DataFrame(data["data"], columns=columns)
print(df.head())

方案2:优化Selenium代码

若必须使用Selenium,优化等待逻辑并直接抓取表格元素的HTML,减少干扰:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

options = Options()
options.add_argument('--headless=new')  # 无头模式提升运行速度

url = "https://www.ssb.no/en/priser-og-prisindekser/konsumpriser/statistikk/konsumprisindeksen"
driver_no = webdriver.Chrome(options=options)

try:
    driver_no.get(url)
    # 等待按钮可点击并执行点击
    btn = WebDriverWait(driver_no, 20).until(
        EC.element_to_be_clickable((By.XPATH, '//*[@id="attachment-table-figure-1"]/button'))
    )
    driver_no.execute_script("arguments[0].click();", btn)
    
    # 等待表格加载完成并获取元素
    table = WebDriverWait(driver_no, 20).until(
        EC.presence_of_element_located((By.ID, 'attachment-table-figure-1-table'))
    )
    # 仅用表格自身的HTML解析,避免整个页面的冗余内容干扰
    df_list = pd.read_html(table.get_attribute('innerHTML'), flavor='html5lib')
    df = df_list[0]
    print(df.head())
finally:
    driver_no.quit()

内容的提问来源于stack exchange,提问作者asuidncsdk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 22:06:32