VertexAI Gemini API安全设置(SafetySettings)无法生效问题排查
VertexAI RAG应用安全过滤配置无效问题排查
我基于VertexAI API开发RAG应用,当查询包含“roubar”一词时,会触发安全过滤器,被标记为HARM_CATEGORY_DANGEROUS_CONTENT(中等风险)而拦截。参照VertexAI官方文档尝试禁用该类别安全过滤或仅拦截高风险等级,但配置后仍无效,以下是代码和报错信息,求排查配置错误。
报错信息
... Candidate: { "index": 0, "finish_reason": "SAFETY", "safety_ratings": [ { "category": "HARM_CATEGORY_HATE_SPEECH", "probability": "NEGLIGIBLE", "probability_score": 0.059326172, "severity": "HARM_SEVERITY_NEGLIGIBLE", "severity_score": 0.10986328 }, { "category": "HARM_CATEGORY_DANGEROUS_CONTENT", "probability": "MEDIUM", "blocked": true, "probability_score": 0.69921875, "severity": "HARM_SEVERITY_MEDIUM", "severity_score": 0.45703125 } ...
现有代码
import os import json from dotenv import load_dotenv import vertexai from vertexai.generative_models import GenerativeModel, GenerationConfig, SafetySetting from promtps import SUBJECT_PROMPT, SYSTEM_PROMPT_3 from utils import VectorDB load_dotenv() # Configs ------------------------------------------------------------------------------ vertexai.init( project=os.environ["VERTEXAI_PROJECT_ID"], location=os.environ["VERTEXAI_REGION"] ) safety_settings = [ SafetySetting( category=SafetySetting.HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT, threshold=SafetySetting.HarmBlockThreshold.BLOCK_NONE, ) ] # RAG flow ----------------------------------------------------------------------------- def subject_identifier(query: str): """ Classify the query subject. """ generation_config_subject = GenerationConfig( response_mime_type="text/plain", temperature=0 ) model_subject = GenerativeModel( model_name=os.environ["VERTEXAI_MODEL"], system_instruction=SUBJECT_PROMPT ) response = model_subject.generate_content(query, generation_config=generation_config_subject) return response.text.strip() def retrieve(subject: str, query:str): """ Retrive documents for the given query. """ vector_db = VectorDB(database_directory=f"data/db/{subject}") chroma_collection = vector_db.get_or_create_chroma_collection() results = chroma_collection.query(query_texts=[query], n_results=4) return results['documents'] def generate(query): """ Embbeds the retrived documents into SYSTEM_PROMPT and generate answers for the given query. """ subject = subject_identifier(query=query) retrieved_documents = retrieve(subject=subject, query=query) generation_config = GenerationConfig( response_mime_type="application/json", temperature=0 ) model = GenerativeModel( model_name=os.environ["VERTEXAI_MODEL"], system_instruction=SYSTEM_PROMPT_3.format(context=retrieved_documents), ) response = model.generate_content( query, generation_config=generation_config, safety_settings=safety_settings ) return json.loads(response.text)
问题排查与修复
你的代码中存在一个关键遗漏:subject_identifier函数调用generate_content时没有传入自定义的safety_settings参数。
RAG流程的第一步是调用subject_identifier对查询进行分类,这一步会先处理包含“roubar”的查询,但此时使用的是VertexAI默认的安全过滤规则,而非你定义的BLOCK_NONE规则,所以直接在这一步就被拦截了,根本没走到后续的generate函数。
修改subject_identifier函数,在generate_content中添加safety_settings参数即可解决:
def subject_identifier(query: str): """ Classify the query subject. """ generation_config_subject = GenerationConfig( response_mime_type="text/plain", temperature=0 ) model_subject = GenerativeModel( model_name=os.environ["VERTEXAI_MODEL"], system_instruction=SUBJECT_PROMPT ) # 添加safety_settings参数 response = model_subject.generate_content( query, generation_config=generation_config_subject, safety_settings=safety_settings ) return response.text.strip()
另外补充说明:如果你希望所有安全类别都使用自定义阈值,建议在safety_settings中明确配置所有HarmCategory,避免部分类别沿用默认规则。比如:
safety_settings = [ SafetySetting( category=SafetySetting.HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT, threshold=SafetySetting.HarmBlockThreshold.BLOCK_NONE, ), SafetySetting( category=SafetySetting.HarmCategory.HARM_CATEGORY_HATE_SPEECH, threshold=SafetySetting.HarmBlockThreshold.BLOCK_NONE, ), SafetySetting( category=SafetySetting.HarmCategory.HARM_CATEGORY_HARASSMENT, threshold=SafetySetting.HarmBlockThreshold.BLOCK_NONE, ), SafetySetting( category=SafetySetting.HarmCategory.HARM_CATEGORY_SEXUALLY_EXPLICIT, threshold=SafetySetting.HarmBlockThreshold.BLOCK_NONE, ) ]
内容的提问来源于stack exchange,提问作者Lucas Miranda de Sena
相关产品推荐
相关产品推荐

