You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup抓取CPI数据并生成有序Pandas DataFrame?

解决方案:通过位置映射提取CPI数据并生成目标DataFrame

既然表格里的<td>元素没有类名或标识,只能靠位置区分字段,那直接按固定索引映射对应字段即可,具体步骤和代码如下:

步骤说明

  1. 先确认每行<td>元素的顺序对应的数据:打印一行的所有<td>文本,明确哪个索引对应年份、Q1-Q4、年度CPI
  2. 遍历表格行时,按索引提取对应数据,同时加上当前国家标识
  3. 收集所有数据后转成Pandas DataFrame,再按国家和年份排序

完整代码示例

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 待抓取的国家列表
target_countries = ["Australia"]
# 存储所有CPI数据的列表
cpi_records = []

for country in target_countries:
    # 替换为实际的澳大利亚CPI历史数据页面URL
    cpi_url = "https://example.com/australia-cpi-history"
    response = requests.get(cpi_url)
    soup = BeautifulSoup(response.text, "html.parser")
    
    # 定位目标表格(根据实际页面调整find的参数,比如id、class)
    cpi_table = soup.find("table")
    # 跳过表头行,遍历数据行
    data_rows = cpi_table.find_all("tr")[1:]
    
    for row in data_rows:
        td_elements = row.find_all("td")
        # 确保当前行有足够的td元素(避免索引越界)
        if len(td_elements) >= 6:
            # 按索引映射字段,根据实际td顺序调整索引!
            record = {
                "country": country,
                "year": td_elements[0].text.strip(),
                "q1": td_elements[1].text.strip(),
                "q2": td_elements[2].text.strip(),
                "q3": td_elements[3].text.strip(),
                "q4": td_elements[4].text.strip(),
                "annual_cpi": td_elements[5].text.strip()
            }
            cpi_records.append(record)

# 转换为DataFrame并排序
cpi_df = pd.DataFrame(cpi_records)
# 按国家升序、年份升序排序
cpi_df = cpi_df.sort_values(by=["country", "year"], ascending=[True, True])

# 可选:转换数据类型(年份转整数,CPI值转浮点数)
cpi_df["year"] = pd.to_numeric(cpi_df["year"], errors="coerce")
for col in ["q1", "q2", "q3", "q4", "annual_cpi"]:
    cpi_df[col] = pd.to_numeric(cpi_df[col].str.replace(r"[^0-9.]", ""), errors="coerce")

print(cpi_df.head())

关键注意事项

  • 调整索引映射:运行前先打印一行td_elements的文本内容(比如print([td.text.strip() for td in td_elements])),确认每个字段对应的索引位置,再修改代码里的索引值
  • 数据清洗:网页上的CPI值可能带百分号、逗号等符号,用str.replace清理后再转成数值类型
  • 异常处理:如果某些行的td数量不足,直接跳过,避免报错;也可以用pd.to_numeric的errors="coerce"参数把无效值转为NaN

内容的提问来源于stack exchange,提问作者srgam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 14:10:50