You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas实现DataFrame文本列以逗号结尾时拆分新增指定行

Pandas处理DataFrame逗号结尾行的实现方案

核心思路

不要用逐行插入的低效写法,先批量处理原行的末尾逗号,再批量构造需要新增的行,通过索引排序保证插入位置正确,性能远高于逐行修改,适合任意数据量的场景。

完整实现代码

import pandas as pd

# 1. 构造初始测试数据
df = pd.DataFrame({
    'Text': ['Alex', 'Smith,', 'Other'],
    'Label': ['name', 'name', '0']
})

# 2. 标记所有Text列以逗号结尾的行
comma_end_flag = df['Text'].str.endswith(',')

# 3. 批量移除符合条件行末尾的逗号
df.loc[comma_end_flag, 'Text'] = df.loc[comma_end_flag, 'Text'].str.rstrip(',')

# 4. 构造需要新增的逗号行,索引设为原行索引+0.5,保证排序后刚好插在对应原行后面
insert_rows = pd.DataFrame(
    {'Text': [','] * comma_end_flag.sum(), 'Label': ['0'] * comma_end_flag.sum()},
    index=df[comma_end_flag].index + 0.5
)

# 5. 拼接数据、排序重置索引得到最终结果
result_df = pd.concat([df, insert_rows]).sort_index().reset_index(drop=True)

结果验证

执行print(result_df)输出如下,完全匹配预期结果:

Text Label
0   Alex  name
1  Smith  name
2      ,     0
3  Other     0

逐行遍历写法(不推荐)

如果数据量极小,也可以用遍历逐行追加的方式实现,逻辑更直观但性能差:

data_buffer = []
for _, row in df.iterrows():
    current_text = row['Text']
    if current_text.endswith(','):
        # 存入去掉末尾逗号的当前行
        data_buffer.append({'Text': current_text.rstrip(','), 'Label': row['Label']})
        # 存入新增的逗号行
        data_buffer.append({'Text': ',', 'Label': '0'})
    else:
        data_buffer.append({'Text': current_text, 'Label': row['Label']})

result_df = pd.DataFrame(data_buffer)

内容的提问来源于stack exchange,提问作者vjugieje

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 08:09:26