在Firebase函数中用Vertex AI分类时,如何绕过拦截将违规文本归为Adult?
问题:Vertex AI拦截不当内容时无法归类为Adult类别
我在Firebase函数中集成了Vertex AI的Gemini模型,用于对用户生成内容进行分类,其中包含Adult类别。但当模型检测到不当文本时,会直接拦截响应并抛出错误,无法按指令归类为Adult:
The response was blocked because the input or response may contain descriptions of violence, sexual themes, or otherwise derogatory content. Please try rephrasing your prompt.
当前实现代码
const textModel = 'gemini-1.5-flash'; const vertexAI = new VertexAI({ project: projectId, location: 'us-central1' }); // Instantiate Gemini models const generativeModel = vertexAI.getGenerativeModel({ model: textModel, }); const textToCategorize = `Categorize this content: "${contentDocumentCategories}". Print only 1 of these categories and don't include anything else: Games, Movies & TV, Music, Celebrities, Humor, Pets & Animals, Food & Drink, Vehicles, Technology, Science, Travel & Transportation, Sports, Online Communities, Books & Literature, Anime & Manga, Comics, Shopping, Nature, Random, Toys, Memes, Business, Interests, People, Lifestyle, Politics, Adult.`; const vertexRequest = { contents: [{ role: 'user', parts: [{ text: textToCategorize }] }], }; const vertexResult = await generativeModel.generateContent(vertexRequest); const vertexResponse = vertexResult.response; // Response if (vertexResponse) console.log('Category: ', vertexResponse.candidates[0].content.parts[0].text);
解决方案
1. 调整模型安全过滤规则
Vertex AI默认的安全过滤器会拦截涉敏内容,需在初始化模型时配置safetySettings,降低对应类别的拦截阈值,或根据业务需求调整过滤强度:
const generativeModel = vertexAI.getGenerativeModel({ model: textModel, safetySettings: [ { category: 'HARM_CATEGORY_SEXUALLY_EXPLICIT', threshold: 'BLOCK_ONLY_HIGH' // 仅拦截极端性内容,轻度内容允许处理 }, { category: 'HARM_CATEGORY_VIOLENCE', threshold: 'BLOCK_ONLY_HIGH' }, { category: 'HARM_CATEGORY_DEROGATORY', threshold: 'BLOCK_ONLY_HIGH' }, { category: 'HARM_CATEGORY_HATE_SPEECH', threshold: 'BLOCK_ONLY_HIGH' }, { category: 'HARM_CATEGORY_HARASSMENT', threshold: 'BLOCK_ONLY_HIGH' } ] });
- 阈值说明:
BLOCK_NONE关闭对应类别过滤(风险较高,谨慎使用);BLOCK_ONLY_HIGH仅拦截极端内容;BLOCK_MEDIUM_AND_ABOVE拦截中等及以上风险内容(默认规则)。
2. 优化提示词明确分类规则
修改提示词,明确告知模型即使内容涉敏,也要返回Adult类别,避免模型产生歧义:
const textToCategorize = `请严格按照以下规则分类内容: 内容:"${contentDocumentCategories}" 可选类别:Games, Movies & TV, Music, Celebrities, Humor, Pets & Animals, Food & Drink, Vehicles, Technology, Science, Travel & Transportation, Sports, Online Communities, Books & Literature, Anime & Manga, Comics, Shopping, Nature, Random, Toys, Memes, Business, Interests, People, Lifestyle, Politics, Adult. 规则:无论内容类型如何,必须从上述列表中选择一个类别返回,仅输出类别名称,不得添加其他内容。若内容涉及暴力、性主题或贬损性内容,统一归类为Adult。`;
3. 添加错误捕获与兜底处理
即使调整安全设置,仍可能有极端内容被拦截,此时通过捕获错误直接返回Adult作为兜底方案:
try { const vertexResult = await generativeModel.generateContent(vertexRequest); const vertexResponse = vertexResult.response; if (vertexResponse && vertexResponse.candidates?.length > 0) { console.log('Category: ', vertexResponse.candidates[0].content.parts[0].text); } else { // 无有效候选响应时返回Adult console.log('Category: Adult'); } } catch (error) { // 捕获拦截类错误,返回Adult if (error.message.includes('The response was blocked')) { console.log('Category: Adult'); } else { // 处理其他类型错误 console.error('分类失败:', error); } }
内容的提问来源于stack exchange,提问作者Miha M
相关产品推荐
相关产品推荐

