如何在Pandas DataFrame中拆分特定前缀的对话列表为两列
Pandas DataFrame对话列表拆分方案
需求说明
需要将DataFrame中dialogue_sentences列的每个单元格列表,按前缀C - :\t和A - :\t拆分到新列dialogue_c和dialogue_a,保留原DataFrame结构(每个单元格对应唯一关联数据)。
输入示例
df_txt['dialogue_sentences'][10] = ['', 'C - :\t hello', 'A - :\t good morning can i speak to you nox', 'C - :\t hello', 'A - :\t hy sien', 'C - :\t hawu', 'A - :\t ok', 'C - :\t yebo ngimi', \"A - :\t hello good morning i am thomas we're calling you from a financial services how are you\", 'C - :\t ngiyaphila', \"A - :\t great to know you Nox we're calling you today in relation to your vehicle the mechanic warranty has been expired i'm sure you are aware of this year\", 'C - :\t no', \"A - :\t it's a mechanical plan on the ford\", 'C - :\t hayi so i sold it', \"A - :\t so you don't have this car anymore the ford\", \"C - :\t ja yes i don't phone it\", \"A - :\t or you don't have\", 'C - :\t ne', 'A - :\t any car', 'C - :\t in', \"A - :\t okay alright no problem i'll make a note for all right\", 'C - :\t thank you', 'A - :\t righto sharp bye', '']
期望输出
df_txt['dialogue_c'][10]:
['hello', 'hello', 'hawu', 'yebo ngimi', 'ngiyaphila', 'no', 'hayi so i sold it', \"ja yes i don't phone it\", 'ne', 'in', 'thank you']
df_txt['dialogue_a'][10]:
['good morning can i speak to you nox', 'hy sien', 'ok', \"hello good morning i am thomas we're calling you from a financial services how are you\", \"great to know you Nox we're calling you today in relation to your vehicle the mechanic warranty has been expired i'm sure you are aware of this year\", \"it's a mechanical plan on the ford\", \"so you don't have this car anymore the ford\", \"or you don't have\", 'any car', \"okay alright no problem i'll make a note for all right\", 'righto sharp bye']
实现方案
步骤1:定义对话拆分函数
针对单个单元格的对话列表,筛选对应前缀的条目,提取并整理对话内容:
import pandas as pd def split_dialogue(sentences): # 提取C的对话:匹配前缀后截取制表符后的内容,过滤空条目 dialogue_c = [s.split('\t')[1].strip() for s in sentences if s.startswith('C - :\t')] # 提取A的对话:逻辑同上 dialogue_a = [s.split('\t')[1].strip() for s in sentences if s.startswith('A - :\t')] # 返回Series用于生成新列 return pd.Series([dialogue_c, dialogue_a], index=['dialogue_c', 'dialogue_a'])
步骤2:应用函数到DataFrame
使用apply方法将拆分函数作用于目标列,直接生成两个新列:
# 生成dialogue_c和dialogue_a列 df_txt[['dialogue_c', 'dialogue_a']] = df_txt['dialogue_sentences'].apply(split_dialogue)
验证结果
执行后可通过以下代码确认输出符合预期:
print(df_txt['dialogue_c'][10]) print(df_txt['dialogue_a'][10])
内容的提问来源于stack exchange,提问作者KateOCM
相关产品推荐
相关产品推荐

