You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas拆分DataFrame中的类别/店铺混合列

解决方案

可以利用Pandas的索引规律或分组操作实现需求,这里提供两种高效的处理方法:

方法一:索引移位法(性能更优)

直接利用两行一组的索引规律,筛选出有消费记录的类别行,再关联对应下一行的店铺名称:

import pandas as pd
import numpy as np

# 示例数据
data = {
    "Category/Shop": ["Supermarkets", "Shop A", "Restaurants", "Shop B", "Clothing", "Shop C"],
    "Spent": [150.0, np.nan, 80.0, np.nan, 200.0, np.nan]
}
df = pd.DataFrame(data)

# 1. 提取有消费记录的类别行(原奇数索引行)
category_rows = df[df["Spent"].notna()].copy()

# 2. 提取对应下一行的店铺名称,按索引步长匹配
category_rows["Shop"] = df["Category/Shop"].iloc[1::2].values

# 3. 重命名列并整理结果结构
category_rows = category_rows.rename(columns={"Category/Shop": "Category"})
result = category_rows[["Category", "Shop", "Spent"]]

# 4. 过滤店铺为空的无效行(若存在)
result = result.dropna(subset=["Shop"])

方法二:分组处理法(逻辑更直观)

将相邻两行分为一组,从每组中提取类别、店铺和消费金额:

# 按两行一组分组(通过索引整除2实现分组)
groups = df.groupby(df.index // 2)

# 从每组提取有效数据
result = groups.apply(lambda group: pd.Series({
    "Category": group["Category/Shop"].iloc[0],
    "Shop": group["Category/Shop"].iloc[1] if len(group) > 1 else np.nan,
    "Spent": group["Spent"].dropna().iloc[0] if not group["Spent"].dropna().empty else np.nan
})).reset_index(drop=True)

# 过滤含空值的无效行
result = result.dropna(subset=["Category", "Shop", "Spent"])

关键说明

  • 移位法适合大数据量场景,性能更高效;分组法逻辑清晰,便于后续调整规则。
  • dropna(subset=["Category", "Shop", "Spent"])会过滤掉类别、店铺或消费金额为空的行,确保结果均为有效数据。

内容的提问来源于stack exchange,提问作者Egor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 19:57:17