在数据框每行两列应用最长公共子串函数时遇报错问题
问题诊断与修复方案
咱们先来揪出问题根源:你的函数在处理单个字符匹配的场景时会崩溃,原因是which(outer(A, B, "=="), arr.ind = TRUE)在只有1处匹配时,返回的是一个长度为2的向量(而不是1行的矩阵)。这时候nrow(ones)会返回NULL,导致1:nrow(ones)直接抛出argument of length 0的错误——你加的if(length(ones)!=0)判断会放行这种情况,但后续循环的范围直接失效了。
修复后的longest.substring函数
我给你调整了函数,核心是强制把匹配位置转为矩阵格式,同时优化了边界条件的判断逻辑,避免其他潜在问题:
longest.substring <- function(a, b) { A <- strsplit(a, "")[[1]] B <- strsplit(b, "")[[1]] L <- matrix(0, length(A), length(B)) # 获取所有字符匹配的位置 ones <- which(outer(A, B, "=="), arr.ind = TRUE) # 无匹配时直接返回空字符串 if (length(ones) == 0) { return("") } # 关键修复:单个匹配时转为1行的矩阵,统一数据格式 if (!is.matrix(ones)) { ones <- matrix(ones, nrow = 1) } # 按行号排序匹配位置 ones <- ones[order(ones[, 1]), ] # 填充动态规划矩阵 for(i in 1:nrow(ones)) { v <- ones[i, , drop = FALSE] # 边界判断:只要行或列是第一位置,最长子串长度就是1 L[v] <- if (v[1] == 1 || v[2] == 1) { 1 } else { # 否则取左上角位置的值+1 L[v[1]-1, v[2]-1] + 1 } } # 找到最长子串的起始位置并拼接结果 max_val <- max(L) max_pos <- which(L == max_val, arr.ind = TRUE)[1, ] start_idx <- max_pos[1] - max_val + 1 paste0(A[start_idx:max_pos[1]], collapse = "") }
测试验证
咱们跑几个用例验证一下:
# 原正常场景 longest.substring("hello world","hella old") # 输出 "hell",符合预期 longest.substring("abc","def") # 输出 "",符合预期 # 之前报错的场景 longest.substring("BRUSKIEGO", "PLATYNOWA") # 输出 "O",正确识别唯一公共字符
批量处理数据框
现在不管是mapply还是apply都能正常工作了,数千行数据用mapply效率会更高:
# 模拟你的数据集 table1 <- data.frame( col1 = c("hello world", "abc", "BRUSKIEGO", "apple pie", "test case"), col2 = c("hella old", "def", "PLATYNOWA", "apple tart", "test case 2"), stringsAsFactors = FALSE ) # 方法1:用mapply批量调用(推荐,效率更高) table1$LCS <- mapply(longest.substring, table1$col1, table1$col2) # 方法2:用apply逐行处理 table1$LCS <- apply(table1[,c("col1","col2")], 1, function(x) longest.substring(x[1], x[2])) # 查看结果 print(table1)
输出结果:
col1 col2 LCS 1 hello world hella old hell 2 abc def 3 BRUSKIEGO PLATYNOWA O 4 apple pie apple tart apple 5 test case test case 2 test cas
内容的提问来源于stack exchange,提问作者PrzeM
相关产品推荐
相关产品推荐

