Tabula GUI与Tabula-py提取PDF表格结果不一致,出现NaN值求助
Tabula GUI与Tabula-py提取PDF表格结果不一致,出现NaN值求助
我最近在处理PDF表格提取的问题,一开始用Tabula的图形界面(GUI)工具,选好要提取的区域后导出的CSV完全符合预期,数据完整没有缺失。但当我导出GUI里的模板,用Tabula-py的read_pdf_with_template方法批量提取时,结果里却出现了不少NaN值,甚至还有数据错位的情况,这可把我难住了。
相关细节
提取区域模板参数
我用的模板是针对PDF第3页的这个区域,参数如下:
[{"page":3,"extraction_method":"guess","x1":56.99806280517571,"x2":537.6866864624022,"y1":239.22816177368162,"y2":737.7751863098144,"width":480.6886236572265,"height":498.54702453613277}]
Tabula GUI的提取效果
GUI导出的表格数据很规整,所有行列都对齐,没有缺失值,完全是我想要的样子。
我用的Python代码
import tabula # 用模板提取PDF表格 df_list = tabula.read_pdf_with_template("Report.pdf", "Report.tabula-template.json") df = df_list[0] print(df)
Tabula-py的提取结果
结果里不仅有NaN,还有数据错位(比如第18行开头的5321明显是和前面的列混在一起了),具体输出如下:
21 dicembre, 24 12:08 135 84 53 21 dicembre, 24 20:27 130 0 21 dicembre, 24 12:53 134.0 82.0 70.0 21 dicembre, 24 21:35 130.0 1 21 dicembre, 24 13:00 136.0 86.0 57.0 21 dicembre, 24 22:56 131.0 2 21 dicembre, 24 14:07 137.0 86.0 65.0 21 dicembre, 24 23:40 135.0 3 21 dicembre, 24 14:15 139.0 89.0 60.0 21 dicembre, 24 23:49 125.0 4 21 dicembre, 24 14:31 132.0 81.0 58.0 21 dicembre, 24 23:57 123.0 5 21 dicembre, 24 15:11 137.0 85.0 60.0 22 dicembre, 24 00:20 121.0 6 21 dicembre, 24 15:19 143.0 89.0 61.0 22 dicembre, 24 00:29 122.0 7 21 dicembre, 24 16:21 124.0 75.0 59.0 22 dicembre, 24 00:37 120.0 8 21 dicembre, 24 16:31 131.0 73.0 58.0 22 dicembre, 24 00:45 123.0 9 21 dicembre, 24 16:40 130.0 77.0 55.0 22 dicembre, 24 00:53 110.0 10 21 dicembre, 24 17:22 136.0 81.0 55.0 22 dicembre, 24 01:34 116.0 11 21 dicembre, 24 17:31 138.0 85.0 58.0 22 dicembre, 24 02:15 125.0 12 21 dicembre, 24 18:12 132.0 76.0 50.0 22 dicembre, 24 02:55 121.0 13 21 dicembre, 24 18:53 133.0 81.0 51.0 22 dicembre, 24 03:03 118.0 14 21 dicembre, 24 19:11 123.0 75.0 50.0 22 dicembre, 24 03:43 119.0 15 21 dicembre, 24 19:20 123.0 76.0 54.0 22 dicembre, 24 04:23 118.0 16 21 dicembre, 24 20:03 135.0 83.0 60.0 NaN NaN NaN 17 NaN NaN NaN NaN 22 dicembre, 24 05:44 125.0 18 5321 dicembre, 24 20:12 131.0 80.0 57.0 22 dicembre, 24 05:53 126.0 76 53.1 0 77.0 57.0 1 78.0 61.0 2 80.0 55.0 3 76.0 53.0 4 71.0 50.0 5 72.0 54.0 6 70.0 49.0 7 70.0 50.0 8 70.0 49.0 9 68.0 49.0 10 69.0 50.0 11 73.0 57.0 12 72.0 51.0 13 71.0 49.0 14 72.0 54.0 15 72.0 52.0 16 NaN NaN 17 78.0 54.0 18 76.0 NaN
我的猜测
我觉得问题可能出在这个PDF表格的结构上——表格的左右两部分没有严格对齐,导致Tabula-py的自动识别逻辑(也就是模板里的guess方法)出了问题,但GUI却能正确处理这种情况。有没有大佬知道怎么让Tabula-py和GUI得到一样的结果呀?
备注:内容来源于stack exchange,提问作者Tosamoon
相关产品推荐
相关产品推荐

