You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas:如何将分组多行数据合并为DataFrame单行?

解决多行文本分组数据合并到DataFrame单行的问题

你需要从多行文本中提取分组数据,将每个Group的动物数据合并到DataFrame的单行中,但当前每个动物单独占一行,目标是得到Group1和Group2各一行的结果。以下是两种高效的解决方法:


方法一:手动按Group收集数据(适合新手理解)

原代码的问题是每次检测到动物行就创建新字典并添加到列表,导致同一Group下的不同动物各占一行。我们可以改为为每个Group维护一个数据字典,收集完该Group的所有动物数据后,再将字典加入结果列表。

import pandas as pd

data = """
Jan 2024
Group1 02/02/2024
dog 10 20
cat 21 32
Group2 05/02/2024
dog 23 45
cat 45 65
owl 24 12
monthly
Admin 02 22
clean 05 32
"""

extract = []
current_group = None
current_data = {}

for line in data.splitlines():
    line = line.strip()
    if not line:  # 跳过空行
        continue
    if 'Group' in line:
        # 保存上一个Group的数据(如果存在)
        if current_group is not None:
            current_data['group'] = current_group
            extract.append(current_data)
        # 初始化新Group的存储字典
        current_group = line.split()[0]
        current_data = {'dog': '', 'cat': '', 'owl': ''}
    elif line.startswith(('dog', 'cat', 'owl')):
        animal, val1, _ = line.split()  # 取第一个数值对应期望输出
        current_data[animal] = val1
    elif line == 'monthly':
        # 处理最后一个Group的数据并终止循环
        if current_group is not None:
            current_data['group'] = current_group
            extract.append(current_data)
        break

df = pd.DataFrame(extract)
df = df[['group', 'dog', 'cat', 'owl']]
print(df)

输出结果:

group dog cat owl
0  Group1  10  21    
1  Group2  23  45  24

方法二:利用pandas的groupby聚合处理

如果你想使用groupby,可以先按原逻辑生成每行数据,再通过分组聚合合并同一Group的行:

import pandas as pd

data = """
Jan 2024
Group1 02/02/2024
dog 10 20
cat 21 32
Group2 05/02/2024
dog 23 45
cat 45 65
owl 24 12
monthly
Admin 02 22
clean 05 32
"""

extract = []
group = None
for line in data.splitlines():
    line = line.strip()
    if not line:
        continue
    if 'Group' in line:
        group = line.split()[0]
    elif line.startswith(('dog', 'cat', 'owl')):
        parts = line.split()
        animal = parts[0]
        val = parts[1]
        # 构造仅当前动物有值的行字典
        row = {'group': group, 'dog': '', 'cat': '', 'owl': ''}
        row[animal] = val
        extract.append(row)

df = pd.DataFrame(extract)
# 按group分组,聚合时取每个列的非空值(每个group下每个动物仅出现一次)
df_grouped = df.groupby('group', as_index=False).agg(
    lambda x: x[x != ''].iloc[0] if any(x != '') else ''
)
print(df_grouped)

输出结果与方法一一致。


内容的提问来源于stack exchange,提问作者MickD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 15:12:59