如何在Pandas DataFrame中将含多值列的行拆分为多行?
解决Pandas多值列拆分多行的KeyError问题
问题场景
现有如下结构的Pandas DataFrame:
key term notes Source 156349471 Aasdasd Bleen 20623750 213740505 dfgdfgdfg Blox 33052911 171645239 rtertertert sdffd 15805072|24361871|28885000 156134219 cvdv dsfsdf 20305092|21259293|21905055|23136149 205936689 ddfg dfsewr 34480604 205947819 xvcbfghf svdst 34480604 213902333 jfghd xcvsd 35020164 156133836 cvbcvb xcvsfg 21907755|30098279 156349486 cvbcvb xcv 24880025 156134727 dfgdfgdfg sdfgdfs 24001450
需求是将Source列中用|分隔的多值拆分为多行,其余列值保持不变,原10行最终应变为16行。
用户尝试的代码触发KeyError:
new_data = [] for _, row in master_df.iterrows(): for src in row['Source'].split('|'): new_data.append([row['key', 'term', 'notes', 'Source'], src]) new_df = pd.DataFrame(new_data, columns=['key', 'term', 'notes', 'Source', 'src']) print(new_df)
报错信息:
KeyError: 'key of type tuple not found and not a MultiIndex'
错误原因
代码中row['key', 'term', 'notes', 'Source']的写法错误:Pandas中选取多列需要用列表索引(row[['key', 'term', 'notes']]),而不是直接传入tuple。你的DataFrame是普通单级列索引,无法识别tuple类型的索引键,因此触发KeyError。
解决方案
方案1:修复循环逻辑
调整多列选取的写法,正确复制原行数据并替换拆分后的Source值:
new_data = [] for _, row in master_df.iterrows(): # 拆分Source列的多值 for src in row['Source'].split('|'): # 获取原行的key、term、notes值,转为列表 base_row = row[['key', 'term', 'notes']].tolist() # 添加拆分后的单个Source值 base_row.append(src) new_data.append(base_row) # 创建新DataFrame,列名与原表一致 new_df = pd.DataFrame(new_data, columns=['key', 'term', 'notes', 'Source']) print(new_df)
方案2:Pandas内置高效方法(推荐)
使用str.split将Source列转为列表,再用explode直接拆分多行,无需手动循环:
# 将Source列的字符串按|拆分为列表 master_df['Source'] = master_df['Source'].str.split('|') # 对Source列执行行拆分,其他列自动复制对应值 new_df = master_df.explode('Source', ignore_index=True) print(new_df)
这个方法更简洁高效,适合处理大数据量,同时自动兼容单值行的情况。
内容的提问来源于stack exchange,提问作者James Petts
相关产品推荐
相关产品推荐

