You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

呼叫中心对话AI信息提取扩展项及Python实现技术问询

呼叫中心对话AI提取方案

一、可额外提取的对话信息

结合电动自行车售后场景,除已规划的提取项外,还可提取以下信息:

  • 客户核心诉求细节:具体故障类型(电池掉电快、刹车异响、控制器故障等)、诉求类型(维修、退换、保修咨询、赔偿等)
  • 坐席服务动作:是否安排上门维修、是否提供远程排查指导、是否申请配件补发、是否转接上级处理等
  • 对话合规性指标:坐席是否主动告知保修范围、是否提醒客户保留凭证、是否符合行业售后话术规范
  • 问题解决状态:对话结束时问题是否当场解决、客户是否接受解决方案、是否需要后续跟进
  • 双方发言占比:客户与坐席各自的发言时长/字数占总对话的比例
  • 关键实体信息:客户提及的电动车型号、购买时间、购买渠道、是否在保修期内、客户所在地区
  • 情绪波动节点:客户情绪从平静转为不满/愤怒的触发点、坐席安抚动作是否有效
  • 信息遗漏情况:坐席是否未询问必要信息(如客户联系方式、地址、故障发生场景)
  • 重复诉求次数:客户重复提及同一问题的次数
  • 竞品提及情况:客户是否提到其他品牌电动车的对比评价

二、Python实现信息提取的具体方案

整体流程

转写文本预处理 → 多维度特征提取 → 结果存储至数据库

所需核心库

nltk/spacy(文本预处理)、transformers(预训练模型用于情感、相似度、主题、实体提取)、sentence-transformers(语义相似度计算)、gensim(主题建模)、sqlite3/pymysql(数据库存储)、re(正则匹配)

分步实现示例

1. 文本预处理:拆分对话角色

将转写文本按发言者拆分,分离坐席与客户的独立发言内容:

import re

# 示例转写文本
transcript = """Bob: 您好,请问有什么可以帮您?
Foo: 我的电动车最近电池掉电特别快,充满电只能骑10公里。
Bob: 请问您的车购买多久了?
Foo: 刚买半年,还在保修期里。
Bob: 好的,我安排明天上门给您检测电池。"""

# 拆分发言者与内容
speaker_pattern = re.compile(r'(\w+): (.*)')
dialogue = []
for line in transcript.split('\n'):
    match = speaker_pattern.match(line.strip())
    if match:
        dialogue.append({
            'speaker': match.group(1),
            'content': match.group(2)
        })

# 分离坐席与客户发言
bob_lines = [item['content'] for item in dialogue if item['speaker'] == 'Bob']
foo_lines = [item['content'] for item in dialogue if item['speaker'] == 'Foo']
bob_full = ' '.join(bob_lines)
foo_full = ' '.join(foo_lines)

2. 已规划提取项的实现

  • 情感倾向:使用DistilBERT预训练模型分析双方情感
from transformers import pipeline

sentiment_analyzer = pipeline('sentiment-analysis')
bob_sentiment = sentiment_analyzer(bob_full)[0]
foo_sentiment = sentiment_analyzer(foo_full)[0]
# 结果示例:{'label': 'POSITIVE', 'score': 0.9876}
  • 填充词统计:自定义填充词列表,统计出现次数
filler_words = ['嗯', '啊', '呃', '那个', '这个', '哦']
def count_fillers(text):
    count = 0
    for word in filler_words:
        count += text.count(word)
    return count

bob_fillers = count_fillers(bob_full)
foo_fillers = count_fillers(foo_full)
  • 回答相关性:计算客户问题与坐席回答的语义相似度
from sentence_transformers import SentenceTransformer, util

model = SentenceTransformer('all-MiniLM-L6-v2')
# 取客户核心问题与坐席对应回答
customer_question = foo_lines[0]
agent_response = bob_lines[-1]

emb1 = model.encode(customer_question, convert_to_tensor=True)
emb2 = model.encode(agent_response, convert_to_tensor=True)
similarity_score = util.cos_sim(emb1, emb2).item()
# 得分范围0-1,越接近1相关性越高
  • 提问数量:统计带疑问词或问号的句子
def count_questions(lines):
    question_count = 0
    question_pattern = re.compile(r'(什么|怎么|为什么|哪|谁|吗?|呢?|\?)')
    for line in lines:
        if question_pattern.search(line):
            question_count +=1
    return question_count

bob_questions = count_questions(bob_lines)
foo_questions = count_questions(foo_lines)
  • 通话时长估算:按平均语速150字/分钟计算
total_words = len(' '.join([item['content'] for item in dialogue]).split())
call_duration = round(total_words / 150, 2)  # 单位:分钟

3. 额外提取项的实现

  • 客户核心诉求提取:用抽取式问答模型提取核心诉求
qa_pipeline = pipeline('question-answering', model='bert-large-uncased-whole-word-masking-finetuned-squad')
context = transcript.replace('\n', ' ')
query = '客户的核心诉求是什么?'
core_demand = qa_pipeline(question=query, context=context)['answer']
  • 关键实体提取:用spaCy提取实体(如时间、产品相关信息)
import spacy

nlp = spacy.load('zh_core_web_sm')  # 中文模型,英文场景用en_core_web_sm
doc = nlp(context)
entities = [(ent.text, ent.label_) for ent in doc.ents]
# 结果示例:[('半年', 'DATE'), ('明天', 'DATE')]

4. 结果存储至数据库

以SQLite为例,创建表并插入分析结果:

import sqlite3

# 连接数据库
conn = sqlite3.connect('call_center.db')
cursor = conn.cursor()

# 创建对话分析表
cursor.execute('''
CREATE TABLE IF NOT EXISTS call_analysis (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    call_id TEXT,
    customer_sentiment TEXT,
    agent_sentiment TEXT,
    customer_fillers INTEGER,
    agent_fillers INTEGER,
    relevance_score REAL,
    customer_questions INTEGER,
    agent_questions INTEGER,
    call_duration REAL,
    core_demand TEXT,
    solve_status TEXT
)
''')

# 插入数据示例
cursor.execute('''
INSERT INTO call_analysis (call_id, customer_sentiment, agent_sentiment, customer_fillers, agent_fillers, relevance_score, customer_questions, agent_questions, call_duration, core_demand, solve_status)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
''', ('CALL_001', foo_sentiment['label'], bob_sentiment['label'], foo_fillers, bob_fillers, similarity_score, foo_questions, bob_questions, call_duration, core_demand, '待上门检测'))

conn.commit()
conn.close()

内容的提问来源于stack exchange,提问作者Faraz Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 11:46:04