You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在基于Azure AI Search的Streamlit应用中显示引用信息

解决方案

核心逻辑

要替换答案中的[doc1]类占位符为真实来源信息,需在检索、提示词构建、结果展示三个环节联动:先从Azure AI Search获取含来源字段的检索结果,再让OpenAI用标记关联内容与来源,最后在Streamlit中替换标记并展示完整来源细节。

具体实现步骤

1. 检索时获取来源字段

修改检索函数,明确指定返回metadata_storage_path及其他需要的字段(如文档标题),同时为每个检索结果生成对应占位符并记录映射关系:

def search_azure_index(query):
    from azure.search.documents import SearchClient
    from azure.core.credentials import AzureKeyCredential

    search_client = SearchClient(
        endpoint=AZURE_SEARCH_ENDPOINT,
        index_name=AZURE_SEARCH_INDEX_NAME,
        credential=AzureKeyCredential(AZURE_SEARCH_KEY)
    )
    # 指定返回需要的字段
    results = search_client.search(
        query=query,
        select=["content", "metadata_storage_path", "metadata_title"],
        top=3
    )

    retrieved_docs = []
    source_mapping = {}
    for idx, res in enumerate(results, 1):
        doc_id = f"doc{idx}"
        retrieved_docs.append(f"[{doc_id}]: {res['content']}")
        # 存储占位符与真实来源的映射
        source_mapping[doc_id] = {
            "title": res.get("metadata_title", "无标题文档"),
            "path": res["metadata_storage_path"]
        }
    return "\n".join(retrieved_docs), source_mapping

2. 构建提示词引导来源标记

在给Azure OpenAI的提示词中,明确要求用生成的占位符标注内容来源:

def build_prompt(retrieved_content, user_query):
    prompt = f"""
    基于以下参考内容回答用户问题:
    {retrieved_content}
    用户问题:{user_query}
    回答要求:
    1. 严格参考给定内容作答,不添加外部信息
    2. 用[{doc_id}]格式的标记标注内容对应的参考来源
    """
    return prompt

3. Streamlit中替换占位符并展示来源

调用接口生成答案后,替换占位符为友好标记,并在答案下方列出所有来源详情:

import streamlit as st
import openai

# 初始化Azure OpenAI
openai.api_type = "azure"
openai.api_base = AZURE_OPENAI_ENDPOINT
openai.api_key = AZURE_OPENAI_KEY
openai.api_version = "2023-03-15-preview"

st.title("文档问答助手")
user_query = st.text_input("请输入你的问题")

if user_query:
    # 检索并获取来源映射
    retrieved_content, source_map = search_azure_index(user_query)
    # 构建提示词
    prompt = build_prompt(retrieved_content, user_query)
    # 调用OpenAI生成答案
    response = openai.ChatCompletion.create(
        engine=AZURE_OPENAI_DEPLOYMENT,
        messages=[{"role": "user", "content": prompt}]
    )
    answer = response.choices[0].message.content

    # 替换答案中的占位符为更易懂的标记
    for doc_id, info in source_map.items():
        answer = answer.replace(f"[{doc_id}]", f"[来源{doc_id[-1]}]")

    # 展示答案
    st.subheader("回答")
    st.write(answer)

    # 展示来源详情
    st.subheader("参考来源")
    for doc_id, info in source_map.items():
        st.write(f"- 来源{doc_id[-1]}: 标题:{info['title']} | 存储路径:{info['path']}")

4. 可选优化

如果metadata_storage_path是Blob的完整URL,可以提取文件名简化显示:

# 在检索函数中处理路径
source_mapping[doc_id] = {
    "title": res.get("metadata_title", "无标题文档"),
    "path": res["metadata_storage_path"].split("/")[-1]  # 提取末尾文件名
}

内容的提问来源于stack exchange,提问作者Tanuj Verma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 06:25:07