如何按句号拆分DataFrame文本列并保留ID生成新DataFrame
按句号拆分DataFrame文本字段并保留ID的解决方案
需求说明
现有包含ID和TEXT字段的DataFrame,需按句号拆分TEXT中的句子,同时保留原始ID生成新的DataFrame。例如文本"I loves cats. I hate snakes"会拆分为两行,对应同一ID。
原始DataFrame
ID TEXT 1 This is a msg. Another msg 2 The weather is hot, the water is cold. My hands are freezing
目标结果
ID TEXT 1 This is a msg 1 Another msg 2 The weather is hot, the water is cold 2 My hands are freezing
原始DataFrame构建代码
import pandas as pd df = pd.DataFrame({'ID':[1,2], 'TEXT':['This is a msg. Another msg', 'The weather is hot, the water is cold. My hands are freezing']})
错误原因
直接使用df['TEXT'].astype(str).split('.')会报错,因为Series对象没有全局的split方法,必须使用Series专属的str.split方法来对每个元素执行字符串拆分操作。
解决方案
方法1:使用str.split + explode(推荐,Pandas 0.25+支持)
这是最简洁直观的方法,explode可以直接将列表类型的字段拆分为多行:
# 按句号拆分文本为列表,再将列表展开为多行 df_result = df.assign(TEXT=df['TEXT'].str.split('.')).explode('TEXT') # 清理拆分后文本两端的空白字符 df_result['TEXT'] = df_result['TEXT'].str.strip() # 过滤掉拆分后可能出现的空字符串(比如原文本末尾带句号的情况) df_result = df_result[df_result['TEXT'] != ''] # 输出结果 print(df_result)
方法2:使用str.split + stack(兼容旧版Pandas)
如果你的Pandas版本低于0.25,不支持explode,可以用stack实现:
# 将ID设为索引,拆分文本为多列,再将列转为行 df_result = df.set_index('ID')['TEXT'].str.split('.', expand=True).stack().reset_index(level=1, drop=True).reset_index(name='TEXT') # 清理空白并过滤空值 df_result['TEXT'] = df_result['TEXT'].str.strip() df_result = df_result[df_result['TEXT'] != ''] print(df_result)
内容的提问来源于stack exchange,提问作者datashout
相关产品推荐
相关产品推荐

