You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R Shiny二元语法词预测应用在Shiny Apps部署异常求助

二元语法Shiny词预测应用部署故障排查

我开发了一款基于二元语法(bi-grams)的Shiny词预测应用,功能是根据输入的首个单词预测第二个单词。该应用在R Studio本地运行完全正常,预测耗时约5秒;但部署到Shiny Apps平台后无法正常工作,预测操作会耗时20秒后触发服务器断开连接。

应用代码

library(shiny)
library(NLP)
library(tibble)
library(tidytext)
library(dplyr)
library(stringr)

ui <- fluidPage(

    titlePanel("Word Prediction with n-Grams by Humberto Renteria - Data Science Capstone"),

    sidebarLayout(
        sidebarPanel(
          textInput("name", "Please enter the word to predict"),
          actionButton("do", "Predict!")
        ),

        mainPanel(
          textOutput("distPlot")
        )
    )
)

server <- function(input, output) {
  
  news_text <- readLines(file("en_US.news.txt", open="r"))
  newsLinesDF <- data_frame(line = 1:length(news_text), text = news_text)
  newsBigrams <- newsLinesDF %>% unnest_tokens(bigram, 
                                               text, token = "ngrams", n = 2)
  
  prediction <- eventReactive(input$do, {
    word_to_start_with <- input$name
    
    last_word <- str_extract(word_to_start_with, "\\b\\w+\\b$")
    
    result <- newsBigrams %>%
      filter(str_detect(bigram, paste0("^", last_word, "\\b"))) %>%
      mutate(second_word = str_extract(bigram, "\\b\\w+\\b")) %>%
      arrange(line) %>%
      slice(1) %>%
      pull(bigram)
    
    return(result)
  })
  
  output$distPlot <- renderText({
    prediction()
  })
}

# Run the application 
shinyApp(ui = ui, server = server)

运行截图

本地R Studio运行状态

本地R Studio运行截图

Shiny Apps部署后状态

Shiny Apps部署后截图

问题排查与优化建议

1. 数据预处理重复执行问题

当前代码将文本读取、二元语法生成逻辑放在server函数顶层,每次用户连接应用都会重复执行这些耗时操作,Shiny Apps服务器资源有限,直接导致超时。

优化方案:
本地提前完成数据预处理并保存结果,部署时直接读取预处理后的文件:

# 本地预处理代码(执行后保存结果)
news_text <- readLines(file("en_US.news.txt", open="r"))
newsLinesDF <- data_frame(line = 1:length(news_text), text = news_text)
newsBigrams <- newsLinesDF %>% unnest_tokens(bigram, text, token = "ngrams", n = 2)
saveRDS(newsBigrams, "newsBigrams.rds")

# 服务器端修改为读取预处理文件
server <- function(input, output) {
  newsBigrams <- readRDS("newsBigrams.rds")
  # 后续预测逻辑保持不变
}

2. 正则表达式过滤效率问题

使用str_detect进行正则匹配的效率较低,可拆分二元语法为结构化字段,直接通过列值匹配:

# 预处理时拆分二元语法
newsBigrams <- newsLinesDF %>% 
  unnest_tokens(bigram, text, token = "ngrams", n = 2) %>%
  separate(bigram, c("first_word", "second_word"), sep = " ")

# 预测时直接匹配first_word
result <- newsBigrams %>%
  filter(first_word == last_word) %>%
  arrange(line) %>%
  slice(1) %>%
  mutate(bigram = paste(first_word, second_word)) %>%
  pull(bigram)

3. 资源适配优化

Shiny Apps免费版资源有限,若数据集过大,可:

  • 对原始文本进行抽样缩减数据量
  • 改用data.table替代dplyr提升查询速度

内容的提问来源于stack exchange,提问作者Humberto R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 19:37:16