You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于File_date与Symbol多列合并两个Pandas DataFrame

基于File_date和Symbol列合并Pandas DataFrame

嘿,我来帮你搞定这两个DataFrame的连接操作!你需要基于File_date和Symbol这两个共同列来合并df_clean和df_allhv_to_date,下面是具体的实现方法和细节:

核心合并方法:pd.merge()

Pandas的merge()函数是处理这种多列连接的最佳选择,你只需要指定连接的键(也就是File_date和Symbol),再根据你的需求选择连接类型即可。

1. 内连接(默认)

只保留两个DataFrame中File_date和Symbol都匹配的行:

import pandas as pd

# 内连接
merged_df = pd.merge(df_clean, df_allhv_to_date, on=['File_date', 'Symbol'])

2. 左连接

保留df_clean的所有行,匹配df_allhv_to_date中对应的行,没有匹配的部分会填充为NaN:

merged_df_left = pd.merge(df_clean, df_allhv_to_date, on=['File_date', 'Symbol'], how='left')

3. 右连接

保留df_allhv_to_date的所有行,匹配df_clean中对应的行:

merged_df_right = pd.merge(df_clean, df_allhv_to_date, on=['File_date', 'Symbol'], how='right')

4. 全外连接

保留两个DataFrame的所有行,没有匹配的部分填充为NaN:

merged_df_outer = pd.merge(df_clean, df_allhv_to_date, on=['File_date', 'Symbol'], how='outer')

注意事项:处理索引问题

从你提供的df_clean索引信息来看,它的索引存在重复值:

df_clean.index
Int64Index([4609, 4611, 4608, 4606, 4603, 4600, 4609, 4607, 4604, 4604, ...
0, 0, 0, 0, 0, 0, 0, 0, 0, 4617], dtype='int64', length=419721)

如果你不需要保留原索引,可以在合并时添加ignore_index=True来生成新的连续索引:

merged_df = pd.merge(df_clean, df_allhv_to_date, on=['File_date', 'Symbol'], ignore_index=True)

或者如果你想保留原索引,可以先重置索引:

df_clean_reset = df_clean.reset_index(drop=False)  # drop=False保留原索引为单独列
merged_df = pd.merge(df_clean_reset, df_allhv_to_date, on=['File_date', 'Symbol'])

补充:你的df_clean示例数据

你提供的df_clean前100行数据如下:

File_date Symbol hv20 hv50 hv100 Date curiv Days Percentile Close Changed
4609 20180423 ZYNE 68 64 64.0 180423.0 65.86 430.0 11.0 10.36 1

内容的提问来源于stack exchange,提问作者sio2bagger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:30:00