You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解析Excel数据并构建可按产品类调用的目标DataFrame?

实现通过df['ProductA']直接访问对应产品类数据的方案

我来帮你搞定这个需求,咱们分两步走:先给每一行数据标记对应的产品类,再实现你想要的直接访问功能。

步骤1:预处理数据,给每行标记产品类

首先读取Excel数据,然后把每个数据行关联到它所属的产品类(比如ProductA),同时移除那些作为分类标题的行(也就是Quantity为NaN的行):

import pandas as pd
import numpy as np

# 读取你的Excel文件
df = pd.read_excel(r'testing.xlsx')

# 1. 创建ProductClass列:仅给标题行(Quantity为NaN)赋值对应的产品类名称
df['ProductClass'] = np.where(df['Quantity'].isna(), df['ItemName'], np.nan)

# 2. 向前填充ProductClass,让标题行下面的所有数据行都继承这个产品类标记
df['ProductClass'] = df['ProductClass'].ffill()

# 3. 过滤掉标题行,只保留有有效Quantity值的数据行
df = df[~df['Quantity'].isna()]

处理后的DataFrame结构大概是这样的(截取部分数据):

ItemNameCategoryQuantityProductClass
AElectronics1.0ProductA
BElectronics2.0ProductA
GHardware7.0ProductB
KSoftware11.0ProductC

步骤2:实现df['ProductA']的直接访问功能

要让你直接通过df['ProductA']获取对应产品类的所有数据,我们可以自定义一个继承自pandas.DataFrame的子类,重写它的__getitem__方法——这样当你输入的键不是列名时,它会自动去查找对应的产品类子集:

class ProductDataFrame(pd.DataFrame):
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        # 按ProductClass分组,生成产品类到子DataFrame的映射字典
        self.product_groups = {name: group for name, group in self.groupby('ProductClass')}
    
    def __getitem__(self, key):
        # 先尝试正常的列访问(比如df['ItemName'])
        try:
            return super().__getitem__(key)
        except KeyError:
            # 如果不是列名,检查是否是产品类名称
            if key in self.product_groups:
                return self.product_groups[key]
            # 如果都找不到,抛出原始的KeyError
            raise KeyError(f"未找到列或产品类 '{key}'")

# 将原DataFrame转换为自定义的ProductDataFrame
df = ProductDataFrame(df)

现在你就可以直接用df['ProductA']获取对应的所有关联数据了:

# 获取ProductA对应的所有数据
print(df['ProductA'])

输出结果会是:

ItemName     Category  Quantity ProductClass
1        A  Electronics       1.0    ProductA
2        B  Electronics       2.0    ProductA
3        C  Electronics       3.0    ProductA
4        D  Electronics       4.0    ProductA
5        E  Electronics       5.0    ProductA
6        F  Electronics       6.0    ProductA

替代方案:无需自定义类,用groupby直接访问

如果你不想自定义类,也可以用pandas自带的groupby功能来实现,只是写法稍有不同:

# 按ProductClass分组
grouped = df.groupby('ProductClass')

# 获取ProductA对应的子集
product_a_data = grouped.get_group('ProductA')

这种方式不需要自定义类,同样能满足需求,只是需要通过grouped.get_group()来访问对应的数据。

内容的提问来源于stack exchange,提问作者viktor_dmitry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 05:47:41