无法编译多个DataFrame数据,报错‘RIC数据类型不被识别’求助
问题解决与优化方案
一、先解决当前报错和代码问题
你遇到的“data type 'RIC' not understood”报错,根源是第一行代码里的dtype='RIC'——pandas的dtype参数需要传入合法的数据类型(比如str、float64),而RIC是你的股票标识列名,不是数据类型,直接写肯定报错。
除此之外你的代码还有几处语法和逻辑错误,修正后的完整代码示例如下(假设read_create_giant_df(file_in)返回的是包含所有股票数据的DataFrame,结构包含Date、RIC、Close等价格列):
import pandas as pd def compile_data(): # 读取所有股票的原始数据,不要乱加dtype参数 all_stocks_df = read_create_giant_df(file_in) # 初始化主DataFrame main_df = pd.DataFrame() # 假设RIC是你的股票代码列表,比如['AAPL', 'MSFT', 'GOOG'] for ric in RIC: # 筛选当前股票的数据 df = all_stocks_df[all_stocks_df['RIC'] == ric].copy() # 设置日期为索引 df.set_index('Date', inplace=True) # 重命名Close列为股票代码 df.rename(columns={'Close': ric}, inplace=True) # 只保留需要的列 df = df[[ric]] # 合并到主DataFrame if main_df.empty: main_df = df else: # 用outer join保留所有日期 main_df = main_df.join(df, how='outer') print(main_df.head()) return main_df compile_data()
二、更简便的合并方法:用pivot一步到位
完全不需要写循环,直接用pandas的pivot方法就能实现你的需求,代码简洁高效:
import pandas as pd def compile_data(): all_stocks_df = read_create_giant_df(file_in) # 直接转置:日期为索引,RIC为列,Close为值 main_df = all_stocks_df.pivot(index='Date', columns='RIC', values='Close') print(main_df.head()) return main_df compile_data()
这个方法会自动把每个RIC对应的Close值放到对应列,日期作为索引,缺失的日期会自动填充NaN,完全符合你的需求。
内容的提问来源于stack exchange,提问作者Catherine Pemblington
相关产品推荐
相关产品推荐

