You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flask+Vosk语音转写至CK Editor的两类技术问题求助

解决方案

问题1:CKEditor原有文本丢失

解决核心是追加识别文本而非直接替换,CKEditor提供多种灵活实现方式:

  • 方式1:获取现有内容后追加到末尾

    // 替换为你的编辑器ID
    const editor = CKEDITOR.instances['editor'];
    // 假设recognizedText是新识别出的文本
    const currentContent = editor.getData();
    // 可根据需求添加空格或换行分隔
    editor.setData(currentContent + ' ' + recognizedText);
    
  • 方式2:插入到当前光标位置(更符合用户输入习惯)

    const editor = CKEDITOR.instances['editor'];
    editor.insertText(' ' + recognizedText);
    

如果后端维护识别历史,也可以让后端返回完整的拼接文本,但前端处理更灵活,避免会话存储的复杂度。

问题2:轮询资源消耗过高

放弃固定间隔轮询,改用事件触发式请求/推送,仅在语音识别期间保持数据交互:

方案A:启停可控的轮询

在开始识别时启动定时器,停止识别时立即清除:

let pollTimer = null;

// 启动语音识别时开启轮询
function startRecognition() {
  pollTimer = setInterval(() => {
    fetch('/api/recognition-result')
      .then(res => res.json())
      .then(data => {
        if (data.text) {
          // 调用之前的编辑器更新函数
          updateEditorContent(data.text);
        }
      });
  }, 600); // 根据识别速度调整间隔,避免过短
}

// 停止语音识别时关闭轮询
function stopRecognition() {
  if (pollTimer) {
    clearInterval(pollTimer);
    pollTimer = null;
  }
}

方案B:WebSocket实时推送(更优)

之前尝试Socket未成功可能是实现细节问题,以下是简化的完整流程:

前端代码

const socket = new WebSocket(`ws://${window.location.host}/ws-recognition`);
let isActive = false;

// 接收后端推送的识别结果
socket.onmessage = (event) => {
  const result = JSON.parse(event.data);
  if (isActive && result.text) {
    updateEditorContent(result.text);
  }
};

// 开始识别
function startRecognition() {
  isActive = true;
  socket.send(JSON.stringify({action: 'start'}));
}

// 停止识别
function stopRecognition() {
  isActive = false;
  socket.send(JSON.stringify({action: 'stop'}));
}

Flask后端代码(使用flask-sock库)

from flask import Flask
from flask_sock import Sock
import json
import vosk
import sounddevice as sd
import sys

app = Flask(__name__)
sock = Sock(app)

# 初始化Vosk模型
model = vosk.Model("path-to-your-model")
recognizer = vosk.KaldiRecognizer(model, 16000)

@sock.route('/ws-recognition')
def handle_recognition_socket(ws):
    while True:
        msg = ws.receive()
        if not msg:
            break
        data = json.loads(msg)
        
        if data['action'] == 'start':
            # 音频捕获回调,实时推送识别结果
            def audio_callback(indata, frames, time, status):
                if status:
                    print(status, file=sys.stderr)
                if recognizer.AcceptWaveform(indata):
                    full_result = json.loads(recognizer.Result())
                    ws.send(json.dumps({'text': full_result['text']}))
                else:
                    partial_result = json.loads(recognizer.PartialResult())
                    ws.send(json.dumps({'text': partial_result['partial']}))
            
            # 启动音频流,直到收到停止信号
            with sd.RawInputStream(samplerate=16000, blocksize=8000, dtype='int16', channels=1, callback=audio_callback):
                while True:
                    stop_msg = ws.receive()
                    if stop_msg and json.loads(stop_msg)['action'] == 'stop':
                        break

CKEditor核心使用要点

CKEditor无需特殊路由,重点在前端交互和后端内容接收:

  • 前端初始化:引入CKEditor资源后,用CKEDITOR.replace('editor-id')初始化编辑器。
  • 内容操作:用getData()获取HTML内容,setData()设置内容,insertText()插入纯文本。
  • 后端接收:表单提交时,编辑器内容会作为对应name的字段传递,后端用request.form.get('editor-name')获取即可。

内容的提问来源于stack exchange,提问作者Musab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 01:53:14