You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将DataFrame中对话按发言者拆分至单独行?

解决方案

要将DataFrame中body列里带HTML标签的多发言者对话,拆分成每行对应一个发言者的Speaker和Transcript列,可按以下步骤操作:

1. 先清洗HTML标签

原始对话包含<p>、<strong>等HTML标签,需先提取纯文本内容。推荐用BeautifulSoup库处理,也可使用正则表达式:

示例代码(BeautifulSoup版)

from bs4 import BeautifulSoup
import pandas as pd

# 模拟用户的原始DataFrame
tempchatdf = pd.DataFrame({
    'body': [
        '<p><strong>Helper:</strong> Hi, I\'m Helper, Virtual Assistant, how can I help you today? Are you inquiring about:eBooksAudiobooks Purchasing Subscriptions Movies etc</p><p><strong>Cx said:</strong> Movies</p>'
    ]
})

# 定义清洗HTML的函数
def clean_html(text):
    soup = BeautifulSoup(text, "html.parser")
    # 用换行符分隔不同发言者的内容
    return soup.get_text(separator="\n")

# 生成清洗后的文本列
tempchatdf['cleaned_body'] = tempchatdf['body'].apply(clean_html)

清洗后cleaned_body列的内容为:

Helper: Hi, I'm Helper, Virtual Assistant, how can I help you today? Are you inquiring about:eBooksAudiobooks Purchasing Subscriptions Movies etc
Cx said: Movies

2. 拆分发言者与对话内容

清洗后的文本每行对应一个发言者,按第一个冒号拆分发言者和对话内容,再将结果展开为多行:

# 按换行符拆分每个发言条目,展开成多行
split_dialogues = tempchatdf['cleaned_body'].str.split('\n', expand=True).stack().reset_index(level=1, drop=True).rename('dialogue')

# 合并回原DataFrame并拆分列
tempchatdf_expanded = tempchatdf.drop(['body', 'cleaned_body'], axis=1).join(split_dialogues)
tempchatdf_expanded[['Speaker', 'Transcript']] = tempchatdf_expanded['dialogue'].str.split(':', n=1, expand=True)

# 去除对话内容的多余空格
tempchatdf_expanded['Transcript'] = tempchatdf_expanded['Transcript'].str.strip()

# 最终结果
final_result = tempchatdf_expanded[['Speaker', 'Transcript']]
print(final_result)

输出结果:

SpeakerTranscript
0HelperHi, I'm Helper, Virtual Assistant, how can I help you today? Are you inquiring about:eBooksAudiobooks Purchasing Subscriptions Movies etc
0Cx saidMovies

3. 替代方案(正则直接提取)

如果不想依赖BeautifulSoup,可直接用正则从原始HTML中匹配发言者和内容:

import re

# 正则模式:匹配<strong>中的发言者,以及后续的对话内容
pattern = re.compile(r'<strong>(.*?):<\/strong>\s*(.*?)(?=<p>|<\/p>|$)', re.DOTALL)

# 提取每个单元格的所有对话对
def extract_dialogues(text):
    matches = pattern.findall(text)
    return [(speaker.strip(), content.strip()) for speaker, content in matches]

# 展开成多行并生成最终DataFrame
dialogues = tempchatdf['body'].apply(extract_dialogues).explode()
final_result = pd.DataFrame(dialogues.tolist(), columns=['Speaker', 'Transcript'], index=dialogues.index)

该方案跳过单独的清洗步骤,直接从HTML中提取目标内容,结果与前一种方法一致。

内容的提问来源于stack exchange,提问作者Hey_jude

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 06:12:43