You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup修正网页表格表头为三行并美化数据?

解决澳大利亚统计局住宅价格指数表格表头整理问题

1. 定位网页表头行

网页里的多级表头通常由多个<tr>标签组成,先通过BeautifulSoup抓取所有表头相关行:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 替换为目标网页URL
target_url = "https://www.abs.gov.au/statistics/economy/price-indexes-and-inflation/residential-property-price-indexes-australia"
resp = requests.get(target_url)
soup = BeautifulSoup(resp.text, "html.parser")

# 抓取表格内前3行表头(根据实际网页结构调整行数)
table = soup.find("table")
header_rows = table.find_all("tr")[:3]

2. 构造三级多级表头

提取每行表头的文本,转换成pandas支持的多级索引(MultiIndex),实现三行表头效果:

# 提取每行表头的文本内容
header_content = []
for row in header_rows:
    cells = row.find_all(["th", "td"])
    row_text = [cell.get_text(strip=True) for cell in cells]
    header_content.append(row_text)

# 生成多级表头
multi_level_header = pd.MultiIndex.from_arrays(header_content)

3. 提取表格数据并匹配表头

跳过表头行抓取数据,将多级表头绑定到DataFrame:

# 抓取数据行(跳过前3行表头)
data_rows = table.find_all("tr")[3:]
table_data = []
for row in data_rows:
    cells = row.find_all(["th", "td"])
    row_data = [cell.get_text(strip=True) for cell in cells]
    table_data.append(row_data)

# 创建带多级表头的DataFrame
df = pd.DataFrame(table_data, columns=multi_level_header)

4. 数据美化与清洗

处理数值格式、空值等,让数据更规范:

# 尝试将列转换为数值类型(跳过非数值列)
for col in df.columns:
    try:
        df[col] = pd.to_numeric(df[col], errors="coerce")
    except Exception:
        continue

# 重置索引(可选,根据需求调整)
df = df.reset_index(drop=True)

提示:如果网页表头的<tr>数量不是3行,可调整代码中切片的数值(比如[:3]和[3:]),确保准确抓取表头和数据行。

内容的提问来源于stack exchange,提问作者ryantl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 09:15:42