You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用GroupBy函数按字符串值拆分单元格并生成分组DataFrame?

解决方案

方法一:使用字符串精确匹配(推荐)

直接检查每行的Conditions列是否包含目标条件,通过正则表达式的单词边界避免子串误匹配:

import pandas as pd

# 初始化示例DataFrame
df = pd.DataFrame({
    "Conditions": ["x y z", "y z", "z", "x"],
    "ID": ["as", "ad", "aw", "am"]
})

# 提取所有唯一的条件值
all_conditions = set()
for cond_str in df["Conditions"]:
    all_conditions.update(cond_str.split())

# 为每个条件生成对应DataFrame
grouped_dfs = {}
for cond in all_conditions:
    # 用单词边界正则确保精确匹配(比如不会把"xy"识别为包含"x")
    grouped_dfs[cond] = df[df["Conditions"].str.contains(rf"\b{cond}\b")]

# 输出结果
print("Dataframe 1(按x分组)")
print(grouped_dfs["x"])
print("\nDataframe 2(按y分组)")
print(grouped_dfs["y"])
print("\nDataframe 3(按z分组)")
print(grouped_dfs["z"])

运行输出:

Dataframe 1(按x分组)
  Conditions  ID
0     x y z  as
3       x  am

Dataframe 2(按y分组)
  Conditions  ID
0     x y z  as
1      y z  ad

Dataframe 3(按z分组)
  Conditions  ID
0     x y z  as
1      y z  ad
2       z  aw

方法二:通过explode展开分组

先将Conditions拆分为列表并展开为多行,再按条件分组去重:

import pandas as pd

df = pd.DataFrame({
    "Conditions": ["x y z", "y z", "z", "x"],
    "ID": ["as", "ad", "aw", "am"]
})

# 拆分条件为列表并展开成多行
df["split_cond"] = df["Conditions"].str.split()
exploded_df = df.explode("split_cond")

# 按展开后的条件分组,去重后得到目标DataFrame
grouped_dfs = {}
for cond, group in exploded_df.groupby("split_cond"):
    grouped_dfs[cond] = group.drop_duplicates(subset=["ID"]).drop(columns=["split_cond"])

两种方法均可实现需求,方法一更简洁高效,方法二更直观,适合复杂拆分场景。

内容的提问来源于stack exchange,提问作者Jeff D

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 09:01:32