如何移除DataFrame表头并转为数据行(Tabula读PDF场景)
问题
我用tabula读取PDF表格,执行代码:
tables = tabula.read_pdf(file, pages="all")
代码能正常运行,tables是DataFrame列表,每个元素对应PDF里的一张表格。但每个DataFrame的第一行被识别成了列名(表头),行索引为0、1、2……
当前DataFrame格式:
Component manufacturer DMNS 0 Component name KL32/OOH8 1 Component type LTE-M/NB-IoT 2 Package markings <pin 1 marker>\ ksdc 99cdjh 3 Date code Not discerned 4 Package type 127-pin land grid array (LGA) 5 Package size 26.00 mm × 10.11 mm × 3.05 mm
期望的DataFrame格式:
0 1 0 Component manufacturer DMNS 1 Component name KL32/OOH8 2 Component type LTE-M/NB-IoT 3 Package markings <pin 1 marker>\ ksdc e99cdjh 4 Date code Not discerned 5 Package type 127-pin land grid array (LGA) 6 Package size 26.00 mm × 10.11 mm × 3.05 mm
如何实现这种格式转换?
解决方案
方法一:读取PDF时直接禁用表头识别
tabula的read_pdf方法提供了header参数,设置为None即可让第一行作为普通数据行,而非列名。修改后的读取代码:
tables = tabula.read_pdf(file, pages="all", header=None)
这样读取的DataFrame会自动用数字(0、1...)作为列名,第一行数据也会保留在DataFrame内,直接符合需求。
方法二:转换已读取的DataFrame
如果已经完成读取,不想重新读取PDF,可以遍历每个DataFrame,将原列名转为第一行数据,再重置列名:
import pandas as pd processed_tables = [] for df in tables: # 将原列名转为一行数据 header_row = pd.DataFrame([df.columns.tolist()], columns=range(df.shape[1])) # 重置原DataFrame的列名为数字索引 df.columns = range(df.shape[1]) # 合并表头行与原数据,重置索引 processed_df = pd.concat([header_row, df], ignore_index=True) processed_tables.append(processed_df)
这段代码兼容任意列数的表格,处理后每个DataFrame都会变成你期望的格式。
内容的提问来源于stack exchange,提问作者spoikayi
相关产品推荐
相关产品推荐

