如何在Pandas聚合操作中合并文本行并插入空格?
解决DataFrame分组聚合文本时自动添加空格的问题
你遇到的问题是因为sum对字符串类型的处理是直接拼接内容,不会自动添加分隔符。要实现合并文本时自动插入空格,只需要把聚合text列的方式从sum换成字符串拼接方法即可,具体有两种简洁的实现方式:
方法一:使用lambda表达式配合str.join
直接在聚合配置中用lambda x: ' '.join(x)替代sum,代码如下:
import pandas as pd # 创建DataFrame df = pd.DataFrame({'id': [101, 101, 102, 102, 102], 'text': ['The government changed', 'the legislation on import control.', 'Politics cannot solve all problems', 'but it should try to do its part.', 'That is the reason why these elections are important.'], 'date': [1990, 1990, 2005, 2005, 2005],}) # 分组聚合,text列用空格拼接 df_new = df.groupby('id').agg({ 'id': 'first', 'text': lambda x: ' '.join(x), 'date': 'first' }).reset_index(drop=True) print(df_new['text'].iloc[0]) # 输出:The government changed the legislation on import control.
方法二:自定义聚合函数(可选)
如果需要后续扩展更复杂的文本处理逻辑,可以单独定义聚合函数:
def join_with_space(text_series): return ' '.join(text_series) df_new = df.groupby('id').agg({ 'id': 'first', 'text': join_with_space, 'date': 'first' }).reset_index(drop=True)
两种方法都能实现文本片段间自动插入空格的效果,推荐第一种lambda写法,简洁直观。
内容的提问来源于stack exchange,提问作者Sarah
相关产品推荐
相关产品推荐

