You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现GPT API返回内容的页面格式化?(LangChain/Flask/Pinecone)

解决GPT返回内容的HTML格式化问题

问题背景

我使用GPT API结合Pinecone、Flask和LangChain实现HTML页面查询功能,当前遇到的问题是:当GPT返回步骤类指令内容时,输出为整段文本,无法自动转换为HTML有序列表,页面展示效果杂乱。

解决方案

有两种可行的处理方式,可根据需求选择:

方法一:让GPT直接返回Markdown格式内容

通过自定义Prompt模板,明确要求GPT返回Markdown有序列表格式的步骤内容,再将Markdown转换为HTML渲染到页面。

1. 修改LangChain的Prompt配置

from langchain.prompts import PromptTemplate

# 自定义Prompt,要求步骤类内容用Markdown有序列表返回
prompt_template = """使用以下上下文回答用户问题。若答案包含操作步骤,请严格使用Markdown有序列表格式输出。

上下文:
{context}

问题: {question}
答案:"""

PROMPT = PromptTemplate(
    template=prompt_template, input_variables=["context", "question"]
)

# 初始化RetrievalQA时传入自定义Prompt
qa_with_source = RetrievalQA.from_chain_type(
    llm=llm, chain_type="stuff", retriever=doc_db.as_retriever(),
    chain_type_kwargs={"prompt": PROMPT}
)

2. 安装Markdown转换库

pip install markdown

3. 修改Flask视图函数,转换Markdown为HTML

import markdown

@app.route("/", methods=["POST", "GET"])
def chat():
    if request.method == "POST":
        user_query = request.form["user_query"]
        message = qa_with_source.run(user_query)
        # 将Markdown转换为HTML
        formatted_message = markdown.markdown(message)
        return render_template(
            "chat.html",
            message=formatted_message)
    else:
        return render_template("chat.html", message=None)

4. 前端模板渲染时允许HTML

在chat.html中使用Jinja2的safe过滤器避免HTML被转义:

{% if message %}
    <div class="chat-response">{{ message|safe }}</div>
{% endif %}

方法二:正则匹配纯文本步骤并转换为HTML

如果不想修改Prompt,可通过正则表达式识别文本中的步骤,手动构建HTML有序列表。

1. 添加文本格式化函数

import re

def convert_steps_to_html(text):
    # 匹配数字开头的步骤(支持跨行内容)
    step_regex = re.compile(r'(\d+\. .+?)(?=\s+\d+\. |\s+Note:|$)', re.DOTALL)
    steps = step_regex.findall(text)
    
    if not steps:
        return text
    
    # 构建HTML有序列表
    html_ol = "<ol>"
    for step in steps:
        html_ol += f"<li>{step.strip()}</li>"
    html_ol += "</ol>"
    
    # 提取开头说明和结尾备注
    intro = re.split(r'\s+\d+\. ', text, maxsplit=1)[0]
    note_match = re.search(r'Note: .+', text, re.DOTALL)
    
    result = intro + html_ol
    if note_match:
        result += f"<p>{note_match.group()}</p>"
    
    return result

2. 在视图函数中调用格式化函数

@app.route("/", methods=["POST", "GET"])
def chat():
    if request.method == "POST":
        user_query = request.form["user_query"]
        message = qa_with_source.run(user_query)
        formatted_message = convert_steps_to_html(message)
        return render_template(
            "chat.html",
            message=formatted_message)
    else:
        return render_template("chat.html", message=None)

3. 前端同样使用safe过滤器渲染

和方法一的前端处理一致,确保HTML能正常解析。

方法对比

  • 方法一:依赖GPT生成结构化Markdown,兼容性好,无需维护正则规则,推荐使用。
  • 方法二:无需修改Prompt,但正则表达式依赖固定的步骤格式,若GPT返回格式变化可能失效,适合特定场景。

内容的提问来源于stack exchange,提问作者William M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 16:23:15