You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

VertexAI Gemini API安全设置(SafetySettings)无法生效问题排查

VertexAI RAG应用安全过滤配置无效问题排查

我基于VertexAI API开发RAG应用,当查询包含“roubar”一词时,会触发安全过滤器,被标记为HARM_CATEGORY_DANGEROUS_CONTENT(中等风险)而拦截。参照VertexAI官方文档尝试禁用该类别安全过滤或仅拦截高风险等级,但配置后仍无效,以下是代码和报错信息,求排查配置错误。

报错信息

...
Candidate:
{
  "index": 0,
  "finish_reason": "SAFETY",
  "safety_ratings": [
    {
      "category": "HARM_CATEGORY_HATE_SPEECH",
      "probability": "NEGLIGIBLE",
      "probability_score": 0.059326172,
      "severity": "HARM_SEVERITY_NEGLIGIBLE",
      "severity_score": 0.10986328
    },
    {
      "category": "HARM_CATEGORY_DANGEROUS_CONTENT",
      "probability": "MEDIUM",
      "blocked": true,
      "probability_score": 0.69921875,
      "severity": "HARM_SEVERITY_MEDIUM",
      "severity_score": 0.45703125
    }
...

现有代码

import os
import json

from dotenv import load_dotenv
import vertexai
from vertexai.generative_models import GenerativeModel, GenerationConfig, SafetySetting

from promtps import SUBJECT_PROMPT, SYSTEM_PROMPT_3
from utils import VectorDB

load_dotenv()

# Configs ------------------------------------------------------------------------------
vertexai.init(
    project=os.environ["VERTEXAI_PROJECT_ID"],
    location=os.environ["VERTEXAI_REGION"]
)

safety_settings = [
    SafetySetting(
        category=SafetySetting.HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT,
        threshold=SafetySetting.HarmBlockThreshold.BLOCK_NONE,
    )
]

# RAG flow -----------------------------------------------------------------------------
def subject_identifier(query: str):
    """
    Classify the query subject.
    """
    generation_config_subject = GenerationConfig(
        response_mime_type="text/plain", 
        temperature=0
    )
    model_subject = GenerativeModel(
        model_name=os.environ["VERTEXAI_MODEL"], 
        system_instruction=SUBJECT_PROMPT
    )
    response = model_subject.generate_content(query, generation_config=generation_config_subject)

    return response.text.strip()


def retrieve(subject: str, query:str):
    """
    Retrive documents for the given query.
    """
    vector_db = VectorDB(database_directory=f"data/db/{subject}")
    chroma_collection = vector_db.get_or_create_chroma_collection()

    results = chroma_collection.query(query_texts=[query], n_results=4)
    
    return results['documents']


def generate(query):
    """
    Embbeds the retrived documents into SYSTEM_PROMPT and generate answers for the given query.
    """
    subject = subject_identifier(query=query)
    retrieved_documents = retrieve(subject=subject, query=query)

    generation_config = GenerationConfig(
        response_mime_type="application/json", 
        temperature=0
    )

    model = GenerativeModel(
        model_name=os.environ["VERTEXAI_MODEL"],
        system_instruction=SYSTEM_PROMPT_3.format(context=retrieved_documents),
    )

    response = model.generate_content(
        query, 
        generation_config=generation_config,
        safety_settings=safety_settings
    )
    
    return json.loads(response.text)

问题排查与修复

你的代码中存在一个关键遗漏:subject_identifier函数调用generate_content时没有传入自定义的safety_settings参数。

RAG流程的第一步是调用subject_identifier对查询进行分类,这一步会先处理包含“roubar”的查询,但此时使用的是VertexAI默认的安全过滤规则,而非你定义的BLOCK_NONE规则,所以直接在这一步就被拦截了,根本没走到后续的generate函数。

修改subject_identifier函数,在generate_content中添加safety_settings参数即可解决:

def subject_identifier(query: str):
    """
    Classify the query subject.
    """
    generation_config_subject = GenerationConfig(
        response_mime_type="text/plain", 
        temperature=0
    )
    model_subject = GenerativeModel(
        model_name=os.environ["VERTEXAI_MODEL"], 
        system_instruction=SUBJECT_PROMPT
    )
    # 添加safety_settings参数
    response = model_subject.generate_content(
        query, 
        generation_config=generation_config_subject,
        safety_settings=safety_settings
    )

    return response.text.strip()

另外补充说明:如果你希望所有安全类别都使用自定义阈值,建议在safety_settings中明确配置所有HarmCategory,避免部分类别沿用默认规则。比如:

safety_settings = [
    SafetySetting(
        category=SafetySetting.HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT,
        threshold=SafetySetting.HarmBlockThreshold.BLOCK_NONE,
    ),
    SafetySetting(
        category=SafetySetting.HarmCategory.HARM_CATEGORY_HATE_SPEECH,
        threshold=SafetySetting.HarmBlockThreshold.BLOCK_NONE,
    ),
    SafetySetting(
        category=SafetySetting.HarmCategory.HARM_CATEGORY_HARASSMENT,
        threshold=SafetySetting.HarmBlockThreshold.BLOCK_NONE,
    ),
    SafetySetting(
        category=SafetySetting.HarmCategory.HARM_CATEGORY_SEXUALLY_EXPLICIT,
        threshold=SafetySetting.HarmBlockThreshold.BLOCK_NONE,
    )
]

内容的提问来源于stack exchange,提问作者Lucas Miranda de Sena

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 06:06:06