如何用Pandas将DataFrame按字段分组后拆分导出为多个TXT文件
问题描述
- 需求:基于DataFrame的
Country和Brand两列做分组,为每一组生成独立的.txt文件 - 原DataFrame结构:

- 原有可正常运行的Excel导出代码:
for (country, brand), group in df_f.groupby(['Country', 'Brand']): group.to_excel(f'{country}_{brand}.xlsx', index=False)
- 原有报错代码的问题点:
- 代码放在了groupby循环外,调用的是完整DataFrame
df_f而非分组后的group - f-string语法错误,引号把
f包裹在了字符串内部 fmt="%d"仅支持纯整数格式,DataFrame中存在字符串类字段时会触发类型错误
正确实现方案
方案1:使用Pandas内置to_csv方法(最简便)
Pandas的to_csv方法原生支持导出为文本格式,无需转换numpy数组,适配各类数据类型:
for (country, brand), group in df_f.groupby(['Country', 'Brand']): # sep参数可自定义分隔符,\t为制表符,如需逗号分隔改为','即可 group.to_csv(f'{country}_{brand}.txt', sep='\t', index=False, encoding='utf-8')
常用参数调整说明:
- 不需要保留表头:添加
header=False参数 - Windows系统打开乱码:将
encoding改为utf-8-sig或gbk - 自定义分隔符:修改
sep参数为对应分隔符号即可
方案2:修正numpy实现写法
如果需要使用numpy实现,调整代码逻辑和参数即可:
for (country, brand), group in df_f.groupby(['Country', 'Brand']): numpy_array = group.to_numpy() # 混合类型数据用%s做通用格式适配,调整f-string写法 np.savetxt(f'{country}_{brand}.txt', numpy_array, fmt = "%s", encoding='utf-8')
内容的提问来源于stack exchange,提问作者MatmataHi
相关产品推荐
相关产品推荐

