You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas read_html抓取网页多表格转DataFrame/CSV格式错乱问题

BLS多表格提取问题修复方案

问题现象

当前已实现基础的BLS网页表格提取功能,单表格页面可正常运行,但处理包含多个统计表格的BLS发布页时,会出现运行报错、提取数据排布混乱、格式错位的问题。

原有代码问题排查

原有实现代码如下:

import urllib
import pandas as pd
from bs4 import BeautifulSoup

def new_func():
    url = input('Please enter the BLS publication that you want to scrape table from:')
    return url

url = new_func()
data = urllib.request.urlopen(url).read()
    
sp = BeautifulSoup(data,'html.parser')

#unwrap multiple table tags in the html
for table in sp.findChildren(attrs={'id': 'regular'}): 
    for c in table.children:
        if c.name in ['tbody', 'thead']:
            c.unwrap()

#dataset creation 
data_pandas = pd.read_html(str(sp), flavor="bs4",thousands=',',decimal='.')
    
#clean dataset
df = pd.concat(data_pandas, axis=0)#to convert lists of pd.read_html to dataframe
    
#export to csv
df.to_csv(input('Specify .csv filename:'))

核心问题点有3个:

  • 解析范围过大:直接把整个页面的BeautifulSoup对象传给pd.read_html,会把页面内导航、脚注、不同统计维度的非目标表格全部读入,无差别拼接必然出现格式混乱
  • 节点修改污染全局DOM树:直接在全局页面对象上执行unwrap()操作,会打破不同表格之间的边界,导致相邻表格的表头、数据行混排,列匹配完全错位
  • 缺少表头校验逻辑:BLS表格自带分类汇总行、表标题、脚注行,依赖pandas自动识别表头很容易出现列数不匹配、列名识别错误的问题

修复方案

调整逻辑为单表格单独处理、单独读取,避免全局DOM污染,同时显式处理表头,修复后参考代码如下:

import urllib
import pandas as pd
from bs4 import BeautifulSoup

def get_target_url():
    url = input('Please enter the BLS publication that you want to scrape table from:')
    return url

url = get_target_url()
data = urllib.request.urlopen(url).read()
sp = BeautifulSoup(data, 'html.parser')

all_dfs = []
# 逐个定位目标表格,不修改全局DOM结构
for idx, table in enumerate(sp.find_all("table", attrs={"id": "regular"})):
    # 单独复制当前表格节点处理多tbody/thead问题,不影响其他表格
    table_clone = BeautifulSoup(str(table), "html.parser")
    for child in table_clone.find_all(["tbody", "thead"]):
        child.unwrap()
    # 读取单个表格,显式指定第一行为表头
    table_df = pd.read_html(
        str(table_clone),
        flavor="bs4",
        thousands=',',
        decimal='.',
        header=0
    )[0]
    # 标记表格序号,方便区分不同表格来源
    table_df["source_table_id"] = idx + 1
    # 去掉全空的无效行
    table_df = table_df.dropna(how="all")
    all_dfs.append(table_df)

# 合并所有有效表格
final_df = pd.concat(all_dfs, axis=0, ignore_index=True)
# 导出csv
final_df.to_csv(input('Specify .csv filename:'), index=False)

额外优化点

  • 把原来命名模糊的new_func改成了语义明确的get_target_url,提升代码可读性
  • 增加了来源表格ID字段,合并后也能区分每行数据来自页面的第几个目标表格
  • 自动过滤全空行,减少后续数据清洗工作量
  • 导出csv时关闭索引写入,避免生成无意义的额外列

内容的提问来源于stack exchange,提问作者Shreyas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 04:51:19