如何用Python将DataFrame中对话按发言者拆分至单独行?
解决方案
要将DataFrame中body列里带HTML标签的多发言者对话,拆分成每行对应一个发言者的Speaker和Transcript列,可按以下步骤操作:
1. 先清洗HTML标签
原始对话包含<p>、<strong>等HTML标签,需先提取纯文本内容。推荐用BeautifulSoup库处理,也可使用正则表达式:
示例代码(BeautifulSoup版)
from bs4 import BeautifulSoup import pandas as pd # 模拟用户的原始DataFrame tempchatdf = pd.DataFrame({ 'body': [ '<p><strong>Helper:</strong> Hi, I\'m Helper, Virtual Assistant, how can I help you today? Are you inquiring about:eBooksAudiobooks Purchasing Subscriptions Movies etc</p><p><strong>Cx said:</strong> Movies</p>' ] }) # 定义清洗HTML的函数 def clean_html(text): soup = BeautifulSoup(text, "html.parser") # 用换行符分隔不同发言者的内容 return soup.get_text(separator="\n") # 生成清洗后的文本列 tempchatdf['cleaned_body'] = tempchatdf['body'].apply(clean_html)
清洗后cleaned_body列的内容为:
Helper: Hi, I'm Helper, Virtual Assistant, how can I help you today? Are you inquiring about:eBooksAudiobooks Purchasing Subscriptions Movies etc Cx said: Movies
2. 拆分发言者与对话内容
清洗后的文本每行对应一个发言者,按第一个冒号拆分发言者和对话内容,再将结果展开为多行:
# 按换行符拆分每个发言条目,展开成多行 split_dialogues = tempchatdf['cleaned_body'].str.split('\n', expand=True).stack().reset_index(level=1, drop=True).rename('dialogue') # 合并回原DataFrame并拆分列 tempchatdf_expanded = tempchatdf.drop(['body', 'cleaned_body'], axis=1).join(split_dialogues) tempchatdf_expanded[['Speaker', 'Transcript']] = tempchatdf_expanded['dialogue'].str.split(':', n=1, expand=True) # 去除对话内容的多余空格 tempchatdf_expanded['Transcript'] = tempchatdf_expanded['Transcript'].str.strip() # 最终结果 final_result = tempchatdf_expanded[['Speaker', 'Transcript']] print(final_result)
输出结果:
| Speaker | Transcript | |
|---|---|---|
| 0 | Helper | Hi, I'm Helper, Virtual Assistant, how can I help you today? Are you inquiring about:eBooksAudiobooks Purchasing Subscriptions Movies etc |
| 0 | Cx said | Movies |
3. 替代方案(正则直接提取)
如果不想依赖BeautifulSoup,可直接用正则从原始HTML中匹配发言者和内容:
import re # 正则模式:匹配<strong>中的发言者,以及后续的对话内容 pattern = re.compile(r'<strong>(.*?):<\/strong>\s*(.*?)(?=<p>|<\/p>|$)', re.DOTALL) # 提取每个单元格的所有对话对 def extract_dialogues(text): matches = pattern.findall(text) return [(speaker.strip(), content.strip()) for speaker, content in matches] # 展开成多行并生成最终DataFrame dialogues = tempchatdf['body'].apply(extract_dialogues).explode() final_result = pd.DataFrame(dialogues.tolist(), columns=['Speaker', 'Transcript'], index=dialogues.index)
该方案跳过单独的清洗步骤,直接从HTML中提取目标内容,结果与前一种方法一致。
内容的提问来源于stack exchange,提问作者Hey_jude
相关产品推荐
相关产品推荐

