You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tabula读PDF得到单列pandas表格,如何按类别拆分为两列?

解决方法

你当前的需求核心是按错位的表头关键词拆分单列数据,分别标记编码类型后合并为两列结构,以下是可直接运行的实现代码:

import pandas as pd

# 此处用你给出的示例数据做演示,实际使用时替换为你读取PDF得到的df即可
data = {'country_code':  [123, 124, 125, 127, 128, 'city_', 'code', 211, 221, 223, 224, 'store_', 'NA', 'code', 321, 3231, 3213, 32123]}
what_i_have = pd.DataFrame(data)

# 1. 定位不同编码类型的分界索引
city_header_idx = what_i_have[what_i_have['country_code'] == 'city_'].index[0]
store_header_idx = what_i_have[what_i_have['country_code'] == 'store_'].index[0]

# 2. 分段提取有效编码,分别标记类型
# 国家编码段:开头到城市表头前,跳过表头行
country_part = pd.DataFrame({
    '编码类型': '国家编码',
    '编码值': what_i_have.loc[:city_header_idx-1, 'country_code']
})
# 城市编码段:城市表头结束后(跳过头2行表头)到门店表头前
city_part = pd.DataFrame({
    '编码类型': '城市编码',
    '编码值': what_i_have.loc[city_header_idx+2 : store_header_idx-1, 'country_code']
})
# 门店编码段:门店表头结束后(跳过头3行表头)到末尾
store_part = pd.DataFrame({
    '编码类型': '门店编码',
    '编码值': what_i_have.loc[store_header_idx+3 :, 'country_code']
})

# 3. 合并所有段并删除空值
result_df = pd.concat([country_part, city_part, store_part], ignore_index=True).dropna()

# 输出最终结果
print(result_df)

适配说明

如果你的实际数据里表头关键词、表头行数和示例不同,只需要调整分界索引的匹配规则,以及对应分段跳过的表头行数即可,编码类型命名也可以按你的需求自定义修改。

内容的提问来源于stack exchange,提问作者NKG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 10:06:04