如何用BeautifulSoup修正网页表格表头为三行并美化数据?
解决澳大利亚统计局住宅价格指数表格表头整理问题
1. 定位网页表头行
网页里的多级表头通常由多个<tr>标签组成,先通过BeautifulSoup抓取所有表头相关行:
import requests from bs4 import BeautifulSoup import pandas as pd # 替换为目标网页URL target_url = "https://www.abs.gov.au/statistics/economy/price-indexes-and-inflation/residential-property-price-indexes-australia" resp = requests.get(target_url) soup = BeautifulSoup(resp.text, "html.parser") # 抓取表格内前3行表头(根据实际网页结构调整行数) table = soup.find("table") header_rows = table.find_all("tr")[:3]
2. 构造三级多级表头
提取每行表头的文本,转换成pandas支持的多级索引(MultiIndex),实现三行表头效果:
# 提取每行表头的文本内容 header_content = [] for row in header_rows: cells = row.find_all(["th", "td"]) row_text = [cell.get_text(strip=True) for cell in cells] header_content.append(row_text) # 生成多级表头 multi_level_header = pd.MultiIndex.from_arrays(header_content)
3. 提取表格数据并匹配表头
跳过表头行抓取数据,将多级表头绑定到DataFrame:
# 抓取数据行(跳过前3行表头) data_rows = table.find_all("tr")[3:] table_data = [] for row in data_rows: cells = row.find_all(["th", "td"]) row_data = [cell.get_text(strip=True) for cell in cells] table_data.append(row_data) # 创建带多级表头的DataFrame df = pd.DataFrame(table_data, columns=multi_level_header)
4. 数据美化与清洗
处理数值格式、空值等,让数据更规范:
# 尝试将列转换为数值类型(跳过非数值列) for col in df.columns: try: df[col] = pd.to_numeric(df[col], errors="coerce") except Exception: continue # 重置索引(可选,根据需求调整) df = df.reset_index(drop=True)
提示:如果网页表头的
<tr>数量不是3行,可调整代码中切片的数值(比如[:3]和[3:]),确保准确抓取表头和数据行。
内容的提问来源于stack exchange,提问作者ryantl
相关产品推荐
相关产品推荐

