Pandas实现DataFrame文本列以逗号结尾时拆分新增指定行
Pandas处理DataFrame逗号结尾行的实现方案
核心思路
不要用逐行插入的低效写法,先批量处理原行的末尾逗号,再批量构造需要新增的行,通过索引排序保证插入位置正确,性能远高于逐行修改,适合任意数据量的场景。
完整实现代码
import pandas as pd # 1. 构造初始测试数据 df = pd.DataFrame({ 'Text': ['Alex', 'Smith,', 'Other'], 'Label': ['name', 'name', '0'] }) # 2. 标记所有Text列以逗号结尾的行 comma_end_flag = df['Text'].str.endswith(',') # 3. 批量移除符合条件行末尾的逗号 df.loc[comma_end_flag, 'Text'] = df.loc[comma_end_flag, 'Text'].str.rstrip(',') # 4. 构造需要新增的逗号行,索引设为原行索引+0.5,保证排序后刚好插在对应原行后面 insert_rows = pd.DataFrame( {'Text': [','] * comma_end_flag.sum(), 'Label': ['0'] * comma_end_flag.sum()}, index=df[comma_end_flag].index + 0.5 ) # 5. 拼接数据、排序重置索引得到最终结果 result_df = pd.concat([df, insert_rows]).sort_index().reset_index(drop=True)
结果验证
执行print(result_df)输出如下,完全匹配预期结果:
Text Label 0 Alex name 1 Smith name 2 , 0 3 Other 0
逐行遍历写法(不推荐)
如果数据量极小,也可以用遍历逐行追加的方式实现,逻辑更直观但性能差:
data_buffer = [] for _, row in df.iterrows(): current_text = row['Text'] if current_text.endswith(','): # 存入去掉末尾逗号的当前行 data_buffer.append({'Text': current_text.rstrip(','), 'Label': row['Label']}) # 存入新增的逗号行 data_buffer.append({'Text': ',', 'Label': '0'}) else: data_buffer.append({'Text': current_text, 'Label': row['Label']}) result_df = pd.DataFrame(data_buffer)
内容的提问来源于stack exchange,提问作者vjugieje
相关产品推荐
相关产品推荐

