使用tabula读PDF得到单列pandas表格,如何按类别拆分为两列?
解决方法
你当前的需求核心是按错位的表头关键词拆分单列数据,分别标记编码类型后合并为两列结构,以下是可直接运行的实现代码:
import pandas as pd # 此处用你给出的示例数据做演示,实际使用时替换为你读取PDF得到的df即可 data = {'country_code': [123, 124, 125, 127, 128, 'city_', 'code', 211, 221, 223, 224, 'store_', 'NA', 'code', 321, 3231, 3213, 32123]} what_i_have = pd.DataFrame(data) # 1. 定位不同编码类型的分界索引 city_header_idx = what_i_have[what_i_have['country_code'] == 'city_'].index[0] store_header_idx = what_i_have[what_i_have['country_code'] == 'store_'].index[0] # 2. 分段提取有效编码,分别标记类型 # 国家编码段:开头到城市表头前,跳过表头行 country_part = pd.DataFrame({ '编码类型': '国家编码', '编码值': what_i_have.loc[:city_header_idx-1, 'country_code'] }) # 城市编码段:城市表头结束后(跳过头2行表头)到门店表头前 city_part = pd.DataFrame({ '编码类型': '城市编码', '编码值': what_i_have.loc[city_header_idx+2 : store_header_idx-1, 'country_code'] }) # 门店编码段:门店表头结束后(跳过头3行表头)到末尾 store_part = pd.DataFrame({ '编码类型': '门店编码', '编码值': what_i_have.loc[store_header_idx+3 :, 'country_code'] }) # 3. 合并所有段并删除空值 result_df = pd.concat([country_part, city_part, store_part], ignore_index=True).dropna() # 输出最终结果 print(result_df)
适配说明
如果你的实际数据里表头关键词、表头行数和示例不同,只需要调整分界索引的匹配规则,以及对应分段跳过的表头行数即可,编码类型命名也可以按你的需求自定义修改。
内容的提问来源于stack exchange,提问作者NKG
相关产品推荐
相关产品推荐

