pandas同年度DataFrame按name合并、跨年度拼接功能实现问询
代码调整方案
原代码存在的问题
- 未按年份对文件分组,直接跨年份横向合并所有文件,不符合需求
- 循环遍历所有文件时重复处理了第一个读入的文件,会产生列名冲突或者重复数据
- 正则匹配提取的年份、变量名是列表格式,未提取有效值
- 未实现不同年度数据的纵向拼接,也没有新增year列标注年份
实现思路
- 先对所有文件路径按提取到的年份分组
- 对每个年份分组内的所有文件,以
Name为键横向合并,整合所有业务列 - 为合并后的年度数据新增
year列赋值为对应年份 - 把所有年度的DataFrame纵向拼接得到最终结果
调整后代码
import pandas as pd import re from collections import defaultdict # 按年份分组存储文件路径 year_file_map = defaultdict(list) for s_path in spatial_paths: # 提取年份,取正则匹配结果的第一个元素转字符串 year = re.findall(r"(\d{4})_spatial_.+\.csv", s_path)[0] year_file_map[year].append(s_path) yearly_dfs = [] # 遍历每个年份处理 for year, file_list in year_file_map.items(): # 读取该年份第一个文件作为初始df base_df = pd.read_csv("directory in google drive" + str(file_list[0])).set_index("Name") # 遍历该年份剩余文件横向合并 for s_path in file_list[1:]: df_new = pd.read_csv("directory in google drive" + str(s_path)).set_index("Name") # how参数可根据需求调整,outer保留所有Name记录,inner只保留共有的Name记录 base_df = pd.merge(base_df, df_new, left_index=True, right_index=True, how="outer") # 新增year列 base_df["year"] = year # 重置索引把Name转回普通列,方便后续拼接 yearly_dfs.append(base_df.reset_index()) # 纵向拼接所有年度数据 final_df = pd.concat(yearly_dfs, ignore_index=True)
内容的提问来源于stack exchange,提问作者FafamKurac123
相关产品推荐
相关产品推荐

