Chrome扩展调用微软语音SDK遇CSP限制:blob媒体加载失败怎么解决
你的场景是Chrome扩展使用microsoft-cognitiveservices-speech-sdk将网页选中文本转换为语音,但部分网站的Content Security Policy(CSP)通过media-src指令限制了媒体资源来源,导致SDK生成的blob URL被拦截,触发类似如下报错:
Refused to load media from 'blob:https://developer.mozilla.org/57fb5680-66f6-49f8-b953-0fb9f2b140df' because it violates the following Content Security Policy directive: "media-src 'self' archive.org videos.cdn.mozilla.net".
以下是几种可行的解决方案:
方案1:直接处理音频流,绕过blob URL生成
微软语音SDK支持在合成完成后直接获取音频二进制数据,无需依赖SpeakerAudioDestination生成blob URL。你可以使用Web Audio API解码并播放这些数据,完全避开页面的media-src限制。
修改后的核心代码示例:
import { getSettings } from "./utils/setting"; let audioContext: AudioContext | null = null; let sourceNode: AudioBufferSourceNode | null = null; export const textToSpeech = async (text: string) => { // 终止当前播放任务 if (sourceNode) { sourceNode.stop(); sourceNode = null; } if (audioContext) { audioContext.close(); audioContext = null; } const voiceName = await getSettings("voice"); return new Promise(async (resolve, reject) => { const speechConfig = window.SpeechSDK.SpeechConfig.fromSubscription(/* 你的订阅参数 */); speechConfig.speechRecognitionLanguage = /* 目标语言,如'zh-CN' */; speechConfig.speechSynthesisVoiceName = voiceName; // 不传入AudioConfig,后续手动处理音频数据 const speechSynthesizer = new window.SpeechSDK.SpeechSynthesizer(speechConfig, null); speechSynthesizer.speakTextAsync( text, async (result) => { if (result.reason === window.SpeechSDK.ResultReason.SynthesizingAudioCompleted) { try { // 初始化AudioContext并解码音频二进制数据 audioContext = new AudioContext(); const audioBuffer = await audioContext.decodeAudioData(result.audioData); // 创建音频源节点并播放 sourceNode = audioContext.createBufferSource(); sourceNode.buffer = audioBuffer; sourceNode.connect(audioContext.destination); sourceNode.start(); // 播放结束后清理资源并resolve sourceNode.onended = () => { resolve(result); speechSynthesizer.close(); }; } catch (decodeError) { reject(decodeError); speechSynthesizer.close(); } } else { reject(new Error(`语音合成失败: ${result.errorDetails}`)); speechSynthesizer.close(); } }, (error) => { console.error(error); reject(error); speechSynthesizer.close(); } ); }); };
方案2:确保代码运行在扩展的隔离世界
Chrome扩展的内容脚本默认运行在隔离世界(Isolated World),这个环境的CSP由扩展自身的manifest.json控制,不受宿主页面的CSP约束。
检查你的manifest.json,确保内容脚本配置正确,避免将代码注入到页面的全局上下文(如通过document.createElement('script')动态插入脚本):
{ "content_scripts": [ { "matches": ["<all_urls>"], // 匹配你需要生效的网站 "js": ["path/to/your/content-script.js"], "world": "ISOLATED" // 默认值,可省略 } ] }
如果之前是通过动态注入脚本到页面全局上下文的方式运行语音合成代码,改为使用隔离世界的内容脚本即可避开页面CSP限制。
方案3:将合成逻辑移至扩展后台/弹出页
扩展的后台脚本(Background Script)、弹出页(Popup)属于扩展自身的上下文,完全不受宿主页面的CSP约束。你可以将语音合成逻辑移至这些位置:
- 内容脚本监听页面选中文本事件,通过
chrome.runtime.sendMessage将文本发送给后台/弹出页 - 后台/弹出页执行语音合成并播放音频
注意:后台脚本默认无法直接播放音频,需要创建一个隐藏的页面(如通过chrome.windows.create创建type: 'popup'且width: 1, height: 1的窗口)来承载音频播放逻辑。
内容的提问来源于stack exchange,提问作者chengfengwang

