解决Pandas读取含重复列名数据时获取重复列的问题
问题描述
- 原始数据表:

- 运行代码:
DTst = gc.open_by_url('https://xxsxx') DTsc = DTst.worksheet('Sales1') DTv = DTsc.get_all_values() DTv = pd.DataFrame.from_records(DTv[1:], columns=DTv[0]) DTUser = DTv.loc[:, ['Items', 'Amount']] DTUser.head()
- 当前输出:

- 问题:
想要的输出是目标图样式,但数据表存在重复列名(见下图),导致现在取出了两列同名的Amount。
解决办法
DataFrame里有重复列名时,用列名直接索引会返回所有同名列,可通过以下三种方式处理:
方式1:按列位置提取
先确定目标列在表中的位置(比如Items是第0列,需要的Amount是第2列),用iloc按位置提取:
# 提取第0列(Items)和第2列(目标Amount) DTUser = DTv.iloc[:, [0, 2]] DTUser.head()
方式2:给重复列名重命名后提取
先给重复列名自动添加后缀区分,再用新列名提取:
# 给重复列名添加序号后缀(比如Amount、Amount_1) DTv.columns = pd.io.parsers.base_parser.ParserBase({'usecols': None})._maybe_dedup_names(DTv.columns) # 根据重命名后的列名提取,比如取第一个Amount或者Amount_1 DTUser = DTv.loc[:, ['Items', 'Amount']] DTUser.head()
方式3:读取数据时直接处理重复列名
在构建DataFrame的步骤就处理重复列名,避免后续操作出错:
# 构建DataFrame时给列名去重 DTv = pd.DataFrame.from_records(DTv[1:], columns=pd.io.parsers.base_parser.ParserBase({'usecols': None})._maybe_dedup_names(DTv[0])) # 再提取需要的列 DTUser = DTv.loc[:, ['Items', 'Amount']] DTUser.head()
内容的提问来源于stack exchange,提问作者khkyt
相关产品推荐
相关产品推荐

