R Shiny二元语法词预测应用在Shiny Apps部署异常求助
二元语法Shiny词预测应用部署故障排查
我开发了一款基于二元语法(bi-grams)的Shiny词预测应用,功能是根据输入的首个单词预测第二个单词。该应用在R Studio本地运行完全正常,预测耗时约5秒;但部署到Shiny Apps平台后无法正常工作,预测操作会耗时20秒后触发服务器断开连接。
应用代码
library(shiny) library(NLP) library(tibble) library(tidytext) library(dplyr) library(stringr) ui <- fluidPage( titlePanel("Word Prediction with n-Grams by Humberto Renteria - Data Science Capstone"), sidebarLayout( sidebarPanel( textInput("name", "Please enter the word to predict"), actionButton("do", "Predict!") ), mainPanel( textOutput("distPlot") ) ) ) server <- function(input, output) { news_text <- readLines(file("en_US.news.txt", open="r")) newsLinesDF <- data_frame(line = 1:length(news_text), text = news_text) newsBigrams <- newsLinesDF %>% unnest_tokens(bigram, text, token = "ngrams", n = 2) prediction <- eventReactive(input$do, { word_to_start_with <- input$name last_word <- str_extract(word_to_start_with, "\\b\\w+\\b$") result <- newsBigrams %>% filter(str_detect(bigram, paste0("^", last_word, "\\b"))) %>% mutate(second_word = str_extract(bigram, "\\b\\w+\\b")) %>% arrange(line) %>% slice(1) %>% pull(bigram) return(result) }) output$distPlot <- renderText({ prediction() }) } # Run the application shinyApp(ui = ui, server = server)
运行截图
本地R Studio运行状态

Shiny Apps部署后状态

问题排查与优化建议
1. 数据预处理重复执行问题
当前代码将文本读取、二元语法生成逻辑放在server函数顶层,每次用户连接应用都会重复执行这些耗时操作,Shiny Apps服务器资源有限,直接导致超时。
优化方案:
本地提前完成数据预处理并保存结果,部署时直接读取预处理后的文件:
# 本地预处理代码(执行后保存结果) news_text <- readLines(file("en_US.news.txt", open="r")) newsLinesDF <- data_frame(line = 1:length(news_text), text = news_text) newsBigrams <- newsLinesDF %>% unnest_tokens(bigram, text, token = "ngrams", n = 2) saveRDS(newsBigrams, "newsBigrams.rds") # 服务器端修改为读取预处理文件 server <- function(input, output) { newsBigrams <- readRDS("newsBigrams.rds") # 后续预测逻辑保持不变 }
2. 正则表达式过滤效率问题
使用str_detect进行正则匹配的效率较低,可拆分二元语法为结构化字段,直接通过列值匹配:
# 预处理时拆分二元语法 newsBigrams <- newsLinesDF %>% unnest_tokens(bigram, text, token = "ngrams", n = 2) %>% separate(bigram, c("first_word", "second_word"), sep = " ") # 预测时直接匹配first_word result <- newsBigrams %>% filter(first_word == last_word) %>% arrange(line) %>% slice(1) %>% mutate(bigram = paste(first_word, second_word)) %>% pull(bigram)
3. 资源适配优化
Shiny Apps免费版资源有限,若数据集过大,可:
- 对原始文本进行抽样缩减数据量
- 改用
data.table替代dplyr提升查询速度
内容的提问来源于stack exchange,提问作者Humberto R
相关产品推荐
相关产品推荐

