在R中使用LLM包批量处理观测值的匹配问题求助
LLM批量调用中Prompt编号匹配异常的排查与解决
需求背景
- 为255个样本×3种提示词的组合(共765次调用)构建带唯一编号的独立Prompt,编号格式如
29C(第29个C类提示词) - 将变量字符串与对应Prompt编号绑定
- 在不触发Token限制的前提下,调用GPT、Claude、Gemini三款LLM
- 将LLM输出存储为DataFrame
- 通过Prompt编号关联输出与原始Prompt
问题现象
GPT和Claude出现Prompt编号识别错误,导致部分观测值被跳过(如示例中的1C、5C)、部分输出重复(如1B、5B各出现两次);Gemini表现正常,无此类异常。
示例输出结果
| prompt_number | variable_food | prompt_type | response |
|---|---|---|---|
| 1A | Bacon, lettuce, tomato, and mayonnaise | basic | No |
| 1B | Bacon, lettuce, tomato, and mayonnaise | ingredients | Yes |
| 1B | Bacon, lettuce, tomato, and mayonnaise | ingredients | No |
| 1C | Bacon, lettuce, tomato, and mayonnaise | NA | NA |
| 2A | Bagel with cream cheese in the middle | basic | No |
| 2B | Bagel with cream cheese in the middle | ingredients | No |
| 2C | Bagel with cream cheese in the middle | named | No |
| 3A | Hot dog | basic | Yes |
| 3B | Hot dog | ingredients | Yes |
| 3C | Hot dog | named | No |
| 4A | Pizza | basic | No |
| 4B | Pizza | ingredients | No |
| 4C | Pizza | named | No |
| 5A | Sandwich Islands | basic | No |
| 5B | Sandwich Islands | ingredients | No |
| 5B | Sandwich Islands | ingredients | No |
| 5C | Sandwich Islands | NA | NA |
模拟场景代码实现
# 加载包与GPT API配置 library(dplyr) library(chatgpt) Sys.setenv(OPENAI_API_KEY = "INSERT API KEY HERE") # 生成数据集与完整Prompt sandwiches <- data.frame( prompt_number = paste0(sort(rep(seq(1:5), 3)), rep(c("A", "B", "C"), 5)), prompt_base = rep( c( "A sandwich is any food item between two pieces of bread. Is the following food item a sandwich? The answer should include two parts: (1) the prompt number and (2) a binary Yes/No response, with no other information. An example of the format of the response text is as follows: \"Prompt number: 29A. Yes,\" where the string 29A and the word \"Yes\" are variable values.", "A sandwich must have a protein, a vegetable, and a condiment. Is the following food item a sandwich? The answer should include two parts: (1) the prompt number and (2) a binary Yes/No response, with no other information. An example of the format of the response text is as follows: \"Prompt number: 29A. Yes,\" where the string 29A and the word \"Yes\" are variable values.", "A sandwich is anything with the word Sandwich in the title. Is the following food item a sandwich? The answer should include two parts, with no other information: (1) the prompt number and (2) a binary Yes/No response. An example of the format of the response text is as follows: \"Prompt number: 29A. Yes,\" where the string 29A and the word \"Yes\" are variable values." ), 5 ), variable_food = sort(rep( c("Pizza", "Hot dog", "Bacon, lettuce, tomato, and mayonnaise", "Sandwich Islands", "Bagel with cream cheese in the middle"), 3 )) ) %>% mutate(full_prompt = paste("Prompt number:", prompt_number, "-", prompt_base, variable_food)) # 定义GPT调用与会话重置函数 ask_and_reset_gpt <- function(x){ Sys.sleep(0.10) ask_and_reset <- list(ask_chatgpt(x), reset_chat_session(x)) results <- ask_and_reset results <- t(results) capture.output(results) } # 初始化结果列表并执行批量调用 results_sandwiches <- list() results_sandwiches <- lapply(sandwiches$full_prompt, ask_and_reset_gpt) # 整理输出结果 results_tidy_sandwiches <- data.frame(unlist(results_sandwiches)) %>% dplyr::rename(output = 1) %>% subset(grepl("Prompt", output)) %>% mutate(output = as.character(output)) %>% separate(output, into = c("discard_1", "output", "discard_2"), sep = "\"") %>% dplyr::select(output) %>% mutate(prompt_number = gsub("Prompt number: ", "", output)) %>% mutate(prompt_type = ifelse(grepl("A", prompt_number), "basic", ifelse(grepl("B", prompt_number), "ingredients", ifelse(grepl("C", prompt_number), "named", NA)))) %>% mutate(response = ifelse(grepl("Yes", prompt_number), "Yes", "No")) %>% mutate(prompt_number = gsub("Yes", "", prompt_number), prompt_number = gsub("No", "", prompt_number)) %>% mutate(prompt_number = trimws(gsub("\\.", "", prompt_number))) # 关联原始Prompt与输出结果 results_and_prompts_sandwiches <- sandwiches %>% left_join(., results_tidy_sandwiches, by = "prompt_number") # 查看结果 results_and_prompts_sandwiches %>% dplyr::select(prompt_number, variable_food, prompt_type, response) %>% head()
问题排查方向
- 提示词工程问题:当前Prompt中编号位置靠后,且格式要求描述冗长,LLM可能忽略或混淆编号信息;示例中的编号与实际编号格式一致但未强化唯一性约束。
- 代码逻辑漏洞:
ask_and_reset_gpt函数中,capture.output(results)可能捕获到非预期输出格式,导致后续解析错误- 结果整理阶段的字符串处理(如
gsub、separate)依赖固定格式,若LLM输出偏离格式则会丢失或错误匹配编号
- LLM包/平台差异:不同LLM对格式指令的遵循度不同,Gemini对结构化输出的支持更稳定;
chatgpt、claudeR包的API封装可能存在会话管理或输出处理的差异。
解决方案建议
1. 提示词优化
- 将Prompt编号前置并强化唯一性,例如:
**必填输出格式:** Prompt编号: {prompt_number} | 答案: Yes/No --- {prompt_base} 食物: {variable_food} - 简化格式要求,用加粗或特殊符号突出必填项,避免冗长描述;移除与当前任务无关的示例编号(如
29A),改用当前实际编号示例。
2. 代码逻辑调整
- 直接在调用时绑定Prompt编号,避免依赖LLM返回编号:
# 修改调用函数,返回编号与输出的绑定结果 ask_gpt_with_id <- function(prompt_id, prompt_text){ Sys.sleep(0.10) response <- ask_chatgpt(prompt_text) reset_chat_session() data.frame(prompt_number = prompt_id, raw_output = response) } # 批量调用时传递编号与Prompt results_sandwiches <- pmap_dfr(list(sandwiches$prompt_number, sandwiches$full_prompt), ask_gpt_with_id) - 优化结果解析逻辑:使用正则表达式精准提取Yes/No,而非依赖编号位置;对解析失败的条目标记为异常,手动处理或重新调用。
3. LLM调用策略优化
- 增加调用间隔(如
Sys.sleep(1)),避免API请求频率过高导致LLM响应质量下降 - 使用LLM的结构化输出功能(如GPT的
response_format="json"),强制返回JSON格式,便于直接解析:# 示例:请求JSON格式输出 ask_chatgpt(prompt_text, model = "gpt-4o", response_format = list(type = "json_schema", json_schema = list( type = "object", properties = list( prompt_number = list(type = "string"), response = list(type = "string", enum = ["Yes", "No"]) ), required = ["prompt_number", "response"] )))
4. 备选方案
- 若R包的API封装存在稳定性问题,可直接使用
httr或curl包调用官方API,自定义请求与响应处理逻辑 - 对于高错误率的Claude,可切换至官方提供的
anthropic包,或调整模型版本(如使用Claude 3 Opus提升格式遵循度)
内容的提问来源于stack exchange,提问作者computer_guy2000
相关产品推荐
相关产品推荐

