You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除DataFrame表头并转为数据行(Tabula读PDF场景)

问题

我用tabula读取PDF表格,执行代码:

tables = tabula.read_pdf(file, pages="all")

代码能正常运行,tables是DataFrame列表,每个元素对应PDF里的一张表格。但每个DataFrame的第一行被识别成了列名(表头),行索引为0、1、2……

当前DataFrame格式:

Component manufacturer               DMNS
0         Component name               KL32/OOH8
1         Component type               LTE-M/NB-IoT
2       Package markings               <pin 1 marker>\ ksdc 99cdjh
3              Date code               Not discerned
4           Package type               127-pin land grid array (LGA)
5           Package size               26.00 mm × 10.11 mm × 3.05 mm

期望的DataFrame格式:

0                                1
0       Component manufacturer           DMNS
1       Component name                   KL32/OOH8
2       Component type                   LTE-M/NB-IoT
3       Package markings                 <pin 1 marker>\ ksdc e99cdjh
4       Date code                        Not discerned
5       Package type                     127-pin land grid array (LGA)
6       Package size                     26.00 mm × 10.11 mm × 3.05 mm

如何实现这种格式转换?


解决方案

方法一:读取PDF时直接禁用表头识别

tabula的read_pdf方法提供了header参数,设置为None即可让第一行作为普通数据行,而非列名。修改后的读取代码:

tables = tabula.read_pdf(file, pages="all", header=None)

这样读取的DataFrame会自动用数字(0、1...)作为列名,第一行数据也会保留在DataFrame内,直接符合需求。

方法二:转换已读取的DataFrame

如果已经完成读取,不想重新读取PDF,可以遍历每个DataFrame,将原列名转为第一行数据,再重置列名:

import pandas as pd

processed_tables = []
for df in tables:
    # 将原列名转为一行数据
    header_row = pd.DataFrame([df.columns.tolist()], columns=range(df.shape[1]))
    # 重置原DataFrame的列名为数字索引
    df.columns = range(df.shape[1])
    # 合并表头行与原数据,重置索引
    processed_df = pd.concat([header_row, df], ignore_index=True)
    processed_tables.append(processed_df)

这段代码兼容任意列数的表格,处理后每个DataFrame都会变成你期望的格式。


内容的提问来源于stack exchange,提问作者spoikayi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 20:18:20