如何将CSV数据重塑为时序结构后通过Fastparquet写入Parquet文件?
实现方案:将CSV时序数据重塑为嵌套结构并存储为Parquet
完全可行!这种将时序数据按country和region维度聚合,把时间相关字段嵌套存储的方式,非常适配Parquet的列式存储特性,而fastparquet库对复杂嵌套类型的读写支持也很完善。下面是具体的实现步骤和代码示例:
1. 准备工作与数据读取
首先确保你已安装必要依赖:
pip install pandas fastparquet
读取CSV数据,并统一时间字段的数值类型:
import pandas as pd import fastparquet # 读取原始CSV df = pd.read_csv('your_input.csv') # 确保year和month为整数类型(若CSV中是字符串格式需转换) df['year'] = df['year'].astype(int) df['month'] = df['month'].astype(int)
2. 重塑为目标嵌套结构
你提到的两种嵌套格式都可以实现,这里分别给出代码:
方式一:元组数组格式
将每个时间点的(year, month, price, volume)打包成元组,再聚合为数组:
def build_tuple_array(group): # 按行打包成元组,转为列表 return list(zip(group['year'], group['month'], group['price'], group['volume'])) # 按country和region分组生成datapoints列 df_tuple_format = df.groupby(['country', 'region'], as_index=False).apply(build_tuple_array, include_groups=False).rename(columns={None: 'datapoints'})
方式二:字典映射格式(更推荐)
以(year, month)为键,{price, volume}为值的字典结构,更便于后续按时间点快速查询:
def build_datapoint_dict(group): # 将year和month设为临时索引,转换为字典格式 group.set_index(['year', 'month'], inplace=True) return group[['price', 'volume']].to_dict('index') # 按country和region分组生成datapoints列 df_dict_format = df.groupby(['country', 'region'], as_index=False).apply(build_datapoint_dict, include_groups=False).rename(columns={None: 'datapoints'})
3. 写入Parquet文件
使用fastparquet将处理后的DataFrame写入Parquet,两种格式都能完美支持:
# 写入字典格式的结果(推荐) df_dict_format.to_parquet('timedata_nested.parquet', engine='fastparquet', index=False) # 也可直接调用fastparquet的write方法 # fastparquet.write('timedata_nested.parquet', df_dict_format, index=False)
4. 验证读取结果
可以读取Parquet文件确认结构是否符合预期:
# 读取Parquet文件 read_back_df = pd.read_parquet('timedata_nested.parquet', engine='fastparquet') # 查看第一条数据的嵌套结构 print(read_back_df.iloc[0]['datapoints'])
补充说明
- Parquet原生支持嵌套数据类型,fastparquet会自动处理字典、数组等复杂结构的序列化与反序列化;
- 这种嵌套结构在后续时序分析中,只需按
country/region过滤,就能直接获取对应维度的所有时序数据,查询效率很高; - 如果数据量较大,fastparquet还支持分块写入、压缩等优化选项,可进一步提升存储和读取性能。
内容的提问来源于stack exchange,提问作者ashic
相关产品推荐
相关产品推荐

