You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中使用LLM包批量处理观测值的匹配问题求助

LLM批量调用中Prompt编号匹配异常的排查与解决

需求背景

  • 为255个样本×3种提示词的组合(共765次调用)构建带唯一编号的独立Prompt,编号格式如29C(第29个C类提示词)
  • 将变量字符串与对应Prompt编号绑定
  • 在不触发Token限制的前提下,调用GPT、Claude、Gemini三款LLM
  • 将LLM输出存储为DataFrame
  • 通过Prompt编号关联输出与原始Prompt

问题现象

GPT和Claude出现Prompt编号识别错误,导致部分观测值被跳过(如示例中的1C、5C)、部分输出重复(如1B、5B各出现两次);Gemini表现正常,无此类异常。

示例输出结果

prompt_numbervariable_foodprompt_typeresponse
1ABacon, lettuce, tomato, and mayonnaisebasicNo
1BBacon, lettuce, tomato, and mayonnaiseingredientsYes
1BBacon, lettuce, tomato, and mayonnaiseingredientsNo
1CBacon, lettuce, tomato, and mayonnaiseNANA
2ABagel with cream cheese in the middlebasicNo
2BBagel with cream cheese in the middleingredientsNo
2CBagel with cream cheese in the middlenamedNo
3AHot dogbasicYes
3BHot dogingredientsYes
3CHot dognamedNo
4APizzabasicNo
4BPizzaingredientsNo
4CPizzanamedNo
5ASandwich IslandsbasicNo
5BSandwich IslandsingredientsNo
5BSandwich IslandsingredientsNo
5CSandwich IslandsNANA

模拟场景代码实现

# 加载包与GPT API配置
library(dplyr)
library(chatgpt)
Sys.setenv(OPENAI_API_KEY = "INSERT API KEY HERE")

# 生成数据集与完整Prompt
sandwiches <- data.frame(
  prompt_number = paste0(sort(rep(seq(1:5), 3)), rep(c("A", "B", "C"), 5)),
  prompt_base = rep(
    c(
      "A sandwich is any food item between two pieces of bread. Is the following food item a sandwich? The answer should include two parts: (1) the prompt number and (2) a binary Yes/No response, with no other information. An example of the format of the response text is as follows: \"Prompt number: 29A. Yes,\" where the string 29A and the word \"Yes\" are variable values.",
      "A sandwich must have a protein, a vegetable, and a condiment. Is the following food item a sandwich? The answer should include two parts: (1) the prompt number and (2) a binary Yes/No response, with no other information. An example of the format of the response text is as follows: \"Prompt number: 29A. Yes,\" where the string 29A and the word \"Yes\" are variable values.",
      "A sandwich is anything with the word Sandwich in the title. Is the following food item a sandwich? The answer should include two parts, with no other information: (1) the prompt number and (2) a binary Yes/No response. An example of the format of the response text is as follows: \"Prompt number: 29A. Yes,\" where the string 29A and the word \"Yes\" are variable values."
    ), 5
  ),
  variable_food = sort(rep(
    c("Pizza", "Hot dog", "Bacon, lettuce, tomato, and mayonnaise", "Sandwich Islands", "Bagel with cream cheese in the middle"), 3
  ))
) %>%
  mutate(full_prompt = paste("Prompt number:", prompt_number, "-", prompt_base, variable_food))

# 定义GPT调用与会话重置函数
ask_and_reset_gpt <- function(x){
  Sys.sleep(0.10)
  ask_and_reset <- list(ask_chatgpt(x), reset_chat_session(x))
  results <- ask_and_reset
  results <- t(results)
  capture.output(results)
}

# 初始化结果列表并执行批量调用
results_sandwiches <- list()
results_sandwiches <- lapply(sandwiches$full_prompt, ask_and_reset_gpt)

# 整理输出结果
results_tidy_sandwiches <- data.frame(unlist(results_sandwiches)) %>%
  dplyr::rename(output = 1) %>%
  subset(grepl("Prompt", output)) %>%
  mutate(output = as.character(output)) %>%
  separate(output, into = c("discard_1", "output", "discard_2"), sep = "\"") %>%
  dplyr::select(output) %>%
  mutate(prompt_number = gsub("Prompt number: ", "", output)) %>%
  mutate(prompt_type = ifelse(grepl("A", prompt_number), "basic",
                              ifelse(grepl("B", prompt_number), "ingredients", 
                                     ifelse(grepl("C", prompt_number), "named", NA)))) %>%
  mutate(response = ifelse(grepl("Yes", prompt_number), "Yes", "No")) %>%
  mutate(prompt_number = gsub("Yes", "", prompt_number),
         prompt_number = gsub("No", "", prompt_number)) %>%
  mutate(prompt_number = trimws(gsub("\\.", "", prompt_number)))

# 关联原始Prompt与输出结果
results_and_prompts_sandwiches <- sandwiches %>%
  left_join(., results_tidy_sandwiches, by = "prompt_number")

# 查看结果
results_and_prompts_sandwiches %>%
  dplyr::select(prompt_number, variable_food, prompt_type, response) %>%
  head()

问题排查方向

  1. 提示词工程问题:当前Prompt中编号位置靠后,且格式要求描述冗长,LLM可能忽略或混淆编号信息;示例中的编号与实际编号格式一致但未强化唯一性约束。
  2. 代码逻辑漏洞:
    • ask_and_reset_gpt函数中,capture.output(results)可能捕获到非预期输出格式,导致后续解析错误
    • 结果整理阶段的字符串处理(如gsub、separate)依赖固定格式,若LLM输出偏离格式则会丢失或错误匹配编号
  3. LLM包/平台差异:不同LLM对格式指令的遵循度不同,Gemini对结构化输出的支持更稳定;chatgpt、claudeR包的API封装可能存在会话管理或输出处理的差异。

解决方案建议

1. 提示词优化

  • 将Prompt编号前置并强化唯一性,例如:
    **必填输出格式:** Prompt编号: {prompt_number} | 答案: Yes/No
    ---
    {prompt_base}
    食物: {variable_food}
    
  • 简化格式要求,用加粗或特殊符号突出必填项,避免冗长描述;移除与当前任务无关的示例编号(如29A),改用当前实际编号示例。

2. 代码逻辑调整

  • 直接在调用时绑定Prompt编号,避免依赖LLM返回编号:
    # 修改调用函数,返回编号与输出的绑定结果
    ask_gpt_with_id <- function(prompt_id, prompt_text){
      Sys.sleep(0.10)
      response <- ask_chatgpt(prompt_text)
      reset_chat_session()
      data.frame(prompt_number = prompt_id, raw_output = response)
    }
    
    # 批量调用时传递编号与Prompt
    results_sandwiches <- pmap_dfr(list(sandwiches$prompt_number, sandwiches$full_prompt), ask_gpt_with_id)
    
  • 优化结果解析逻辑:使用正则表达式精准提取Yes/No,而非依赖编号位置;对解析失败的条目标记为异常,手动处理或重新调用。

3. LLM调用策略优化

  • 增加调用间隔(如Sys.sleep(1)),避免API请求频率过高导致LLM响应质量下降
  • 使用LLM的结构化输出功能(如GPT的response_format="json"),强制返回JSON格式,便于直接解析:
    # 示例:请求JSON格式输出
    ask_chatgpt(prompt_text, model = "gpt-4o", response_format = list(type = "json_schema", json_schema = list(
      type = "object",
      properties = list(
        prompt_number = list(type = "string"),
        response = list(type = "string", enum = ["Yes", "No"])
      ),
      required = ["prompt_number", "response"]
    )))
    

4. 备选方案

  • 若R包的API封装存在稳定性问题,可直接使用httr或curl包调用官方API,自定义请求与响应处理逻辑
  • 对于高错误率的Claude,可切换至官方提供的anthropic包,或调整模型版本(如使用Claude 3 Opus提升格式遵循度)

内容的提问来源于stack exchange,提问作者computer_guy2000

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 13:27:07