呼叫中心对话AI信息提取扩展项及Python实现技术问询
呼叫中心对话AI提取方案
一、可额外提取的对话信息
结合电动自行车售后场景,除已规划的提取项外,还可提取以下信息:
- 客户核心诉求细节:具体故障类型(电池掉电快、刹车异响、控制器故障等)、诉求类型(维修、退换、保修咨询、赔偿等)
- 坐席服务动作:是否安排上门维修、是否提供远程排查指导、是否申请配件补发、是否转接上级处理等
- 对话合规性指标:坐席是否主动告知保修范围、是否提醒客户保留凭证、是否符合行业售后话术规范
- 问题解决状态:对话结束时问题是否当场解决、客户是否接受解决方案、是否需要后续跟进
- 双方发言占比:客户与坐席各自的发言时长/字数占总对话的比例
- 关键实体信息:客户提及的电动车型号、购买时间、购买渠道、是否在保修期内、客户所在地区
- 情绪波动节点:客户情绪从平静转为不满/愤怒的触发点、坐席安抚动作是否有效
- 信息遗漏情况:坐席是否未询问必要信息(如客户联系方式、地址、故障发生场景)
- 重复诉求次数:客户重复提及同一问题的次数
- 竞品提及情况:客户是否提到其他品牌电动车的对比评价
二、Python实现信息提取的具体方案
整体流程
转写文本预处理 → 多维度特征提取 → 结果存储至数据库
所需核心库
nltk/spacy(文本预处理)、transformers(预训练模型用于情感、相似度、主题、实体提取)、sentence-transformers(语义相似度计算)、gensim(主题建模)、sqlite3/pymysql(数据库存储)、re(正则匹配)
分步实现示例
1. 文本预处理:拆分对话角色
将转写文本按发言者拆分,分离坐席与客户的独立发言内容:
import re # 示例转写文本 transcript = """Bob: 您好,请问有什么可以帮您? Foo: 我的电动车最近电池掉电特别快,充满电只能骑10公里。 Bob: 请问您的车购买多久了? Foo: 刚买半年,还在保修期里。 Bob: 好的,我安排明天上门给您检测电池。""" # 拆分发言者与内容 speaker_pattern = re.compile(r'(\w+): (.*)') dialogue = [] for line in transcript.split('\n'): match = speaker_pattern.match(line.strip()) if match: dialogue.append({ 'speaker': match.group(1), 'content': match.group(2) }) # 分离坐席与客户发言 bob_lines = [item['content'] for item in dialogue if item['speaker'] == 'Bob'] foo_lines = [item['content'] for item in dialogue if item['speaker'] == 'Foo'] bob_full = ' '.join(bob_lines) foo_full = ' '.join(foo_lines)
2. 已规划提取项的实现
- 情感倾向:使用DistilBERT预训练模型分析双方情感
from transformers import pipeline sentiment_analyzer = pipeline('sentiment-analysis') bob_sentiment = sentiment_analyzer(bob_full)[0] foo_sentiment = sentiment_analyzer(foo_full)[0] # 结果示例:{'label': 'POSITIVE', 'score': 0.9876}
- 填充词统计:自定义填充词列表,统计出现次数
filler_words = ['嗯', '啊', '呃', '那个', '这个', '哦'] def count_fillers(text): count = 0 for word in filler_words: count += text.count(word) return count bob_fillers = count_fillers(bob_full) foo_fillers = count_fillers(foo_full)
- 回答相关性:计算客户问题与坐席回答的语义相似度
from sentence_transformers import SentenceTransformer, util model = SentenceTransformer('all-MiniLM-L6-v2') # 取客户核心问题与坐席对应回答 customer_question = foo_lines[0] agent_response = bob_lines[-1] emb1 = model.encode(customer_question, convert_to_tensor=True) emb2 = model.encode(agent_response, convert_to_tensor=True) similarity_score = util.cos_sim(emb1, emb2).item() # 得分范围0-1,越接近1相关性越高
- 提问数量:统计带疑问词或问号的句子
def count_questions(lines): question_count = 0 question_pattern = re.compile(r'(什么|怎么|为什么|哪|谁|吗?|呢?|\?)') for line in lines: if question_pattern.search(line): question_count +=1 return question_count bob_questions = count_questions(bob_lines) foo_questions = count_questions(foo_lines)
- 通话时长估算:按平均语速150字/分钟计算
total_words = len(' '.join([item['content'] for item in dialogue]).split()) call_duration = round(total_words / 150, 2) # 单位:分钟
3. 额外提取项的实现
- 客户核心诉求提取:用抽取式问答模型提取核心诉求
qa_pipeline = pipeline('question-answering', model='bert-large-uncased-whole-word-masking-finetuned-squad') context = transcript.replace('\n', ' ') query = '客户的核心诉求是什么?' core_demand = qa_pipeline(question=query, context=context)['answer']
- 关键实体提取:用spaCy提取实体(如时间、产品相关信息)
import spacy nlp = spacy.load('zh_core_web_sm') # 中文模型,英文场景用en_core_web_sm doc = nlp(context) entities = [(ent.text, ent.label_) for ent in doc.ents] # 结果示例:[('半年', 'DATE'), ('明天', 'DATE')]
4. 结果存储至数据库
以SQLite为例,创建表并插入分析结果:
import sqlite3 # 连接数据库 conn = sqlite3.connect('call_center.db') cursor = conn.cursor() # 创建对话分析表 cursor.execute(''' CREATE TABLE IF NOT EXISTS call_analysis ( id INTEGER PRIMARY KEY AUTOINCREMENT, call_id TEXT, customer_sentiment TEXT, agent_sentiment TEXT, customer_fillers INTEGER, agent_fillers INTEGER, relevance_score REAL, customer_questions INTEGER, agent_questions INTEGER, call_duration REAL, core_demand TEXT, solve_status TEXT ) ''') # 插入数据示例 cursor.execute(''' INSERT INTO call_analysis (call_id, customer_sentiment, agent_sentiment, customer_fillers, agent_fillers, relevance_score, customer_questions, agent_questions, call_duration, core_demand, solve_status) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) ''', ('CALL_001', foo_sentiment['label'], bob_sentiment['label'], foo_fillers, bob_fillers, similarity_score, foo_questions, bob_questions, call_duration, core_demand, '待上门检测')) conn.commit() conn.close()
内容的提问来源于stack exchange,提问作者Faraz Ahmed
相关产品推荐
相关产品推荐

