通过OCR提取图片数据并设置列名时遇R语言报错求助
问题:OCR提取图片数据转DataFrame时出现列数不匹配错误
问题背景
我使用以下R代码提取图片中的数据,在尝试设置列名时触发错误:
library(magick) library(tesseract) team_img <- image_read("C:\\Users\\macosta\\Downloads\\Capture1.PNG") %>% image_resize('2000x') %>% image_convert(type = 'Grayscale') %>% image_trim(fuzz = 40) %>% image_write(format = 'png', density = '300x300') %>% tesseract::ocr() %>% strsplit('\n') %>% getElement(1) %>% `[`(-1) %>% {read.table(text = .)} %>% # 此处报错 setNames(c('Number', 'Weight', 'Specie', 'long', 'lat', 'altitude', 'Date', 'Others'))
错误信息
Error in scan(file = file, what = what, sep = sep, quote = quote, dec = dec, :
line 2 did not have 9 elements
OCR提取的文本数据
<- team_img # [1] "" # [2] "3057 4.2 vulgaris 90234\" W 14237N 1550 m 20/1/95" # [3] "3058 10.4 lunatus 90234°W 14237N 1550 m 20/1/95" # [4] "3059 12.0 lunatus 90246’W 14238’N 1700 m 20/1/95" # [5] "3060 4.7 ~~ vulgaris 90247'W 14242’°N 1330 m 20/1/95" # [6] "3061 4.7 vulgaris 90247 W 14243’N 1350 m 20/1/95" # [7] "3062 10.8 coccineus 90951,W 14933N 2080m 20/1/95" # [8] "3063 6.8 __lunatus 90248°W —-14221’N 520 m 21/1/95" # [9] "3064 5.8 lunatus 91211\"W 14°24°N 230 m 21/1/95" # [10] "3065 7.0 lunatus 91208’W 14231’N 750 m 21/1/95" # [11] "3066 4.2 vulgaris 91207 W 14236°N 1330 m 21/1/95" # [12] "3067 23.3 polyanth. 91207°W 14238’N 1630 m 21/1/95" # [13] "3068 10.4 coccineus 91222’W 15221’N 2070 m 22/1/95" # [14] "3069 8.0 coccineus 91222’W 15221’N 2070 m 22/1/95" # [15] "3070. --- leptosta. 91221’W 15221’N 1910 m 22/1/95" # [16] "3071 7.8 lunatus 91254°W 15241’N 590 m 23/1/95" # [17] "3072 5.0 Cnido.sp. 91254°W 15241’N 600 m 23/1/95" # [18] "3073 0.6 _leptosta. 91248°W 15241’N 1000 m 23/1/95" # [19] "3074 3.9 vulgaris 91247'W 15240°N 1180 m 23/1/95" # [20] "3075 4.0 vulgaris 91942\"W 15939°N 1540 m 23/1/95" # [21] "3076 6.4 coccineus 91231’W 14249’°N 2370 m 24/1/95" # [22] "3077 9.5 coccineus 91229°W 14248°N 2070 m 24/1/95" # [23] "3078 64 coccineus 91229°W 14247N 2050 m 24/1/95" # [24] "3079 11.8 vulgaris 91229’W 14247'N 2050 m 24/1/95" # [25] "3080 10.5 lunatus 91230°W 14246’N 1750 m 24/1/95" # [26] "3081 6.9 vulgaris 91231’°W 14245’N 1730 m 24/1/95"
数据图片

目标
将提取的数据整理为标准DataFrame
解决方案
问题根源
报错核心是OCR提取的文本格式不一致:
- 部分行的经纬度被拆分为多段(如第2行
90234" W是两个元素,正常行是90234°W单个元素) - 存在冗余符号(
~~、__、---) - 部分数值格式错误(如第15行
3070. ---、第23行64缺少小数点)
修正代码
library(magick) library(tesseract) library(stringr) library(purrr) # 图片处理与OCR提取 team_img <- image_read("C:\\Users\\macosta\\Downloads\\Capture1.PNG") %>% image_resize('2000x') %>% image_convert(type = 'Grayscale') %>% image_trim(fuzz = 40) %>% image_write(format = 'png', density = '300x300') %>% tesseract::ocr() %>% strsplit('\n') %>% getElement(1) %>% `[`(-1) %>% # 移除空行 # 统一文本格式 str_replace_all("~~|__|---", "") %>% # 删除冗余符号 str_replace_all("(\\d+)[\\s°’'\"](W|N)", "\\1\\2") %>% # 合并经纬度数字与方向 str_replace_all("(\\d+)\\s*(m)", "\\1\\2") %>% # 合并海拔数字与单位 str_replace_all("(\\d+)\\.(\\d+)\\.(\\d+)", "\\1.\\2") %>% # 修正编号多余小数点 str_replace_all("(\\d+)(?=[A-Za-z.])", "\\1 ") %>% # 给无空格的数字补空格 keep(~str_squish(.) != "") # 过滤清理后仍为空的行 # 正则匹配拆分列并转换为DataFrame df <- str_match_all(team_img, "^(\\d+)\\s+(\\d*\\.?\\d*)\\s+([A-Za-z.]+)\\s+([\\d°’'\"W]+)\\s+([\\d°’'\"N]+)\\s+(\\d+m)\\s+(\\d+/\\d+/\\d+)\\s*(.*)$") %>% map_df(function(x) { data.frame( Number = x[,2], Weight = as.numeric(x[,3]), Specie = x[,4], long = x[,5], lat = x[,6], altitude = x[,7], Date = x[,8], Others = x[,9], stringsAsFactors = FALSE ) }) # 查看结果 head(df)
代码说明
str_replace_all系列函数:统一OCR文本格式,解决拆分时的列数不一致问题str_match_all:用正则表达式精确匹配8列结构,确保每行都能正确拆分map_df:将拆分后的列表直接转换为DataFrame
内容的提问来源于stack exchange,提问作者Miguel Angel Acosta Chinchilla
相关产品推荐
相关产品推荐

