You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对比两个Pandas DataFrame找出缺失的城市行?

找出DataFrame中缺失的城市名称

方法1:集合差集快速定位缺失值

直接提取两个DataFrame的cities列转为集合,通过差集运算快速得到缺失的城市:

# 找出df2中没有、但df1有的城市
missing_in_df2 = set(df1['cities']) - set(df2['cities'])
print("df2缺失的城市:", missing_in_df2)

# 反过来,找出df1中没有、但df2有的城市(按需使用)
missing_in_df1 = set(df2['cities']) - set(df1['cities'])
print("df1缺失的城市:", missing_in_df1)

方法2:用isin()筛选缺失的完整行

如果需要获取缺失城市对应的整行数据,可用isin()配合取反操作:

# 获取df1存在、但df2不存在的行
missing_rows_in_df2 = df1[~df1['cities'].isin(df2['cities'])]
print("df2缺失的行:")
print(missing_rows_in_df2)

# 获取df2存在、但df1不存在的行(按需使用)
missing_rows_in_df1 = df2[~df2['cities'].isin(df1['cities'])]
print("df1缺失的行:")
print(missing_rows_in_df1)

方法3:全外连接标记差异

通过merge()做全外连接并标记匹配情况,直观展示两边的差异:

merged = df1.merge(df2, on='cities', how='outer', indicator=True)
# 筛选仅在df1存在的城市
only_in_df1 = merged[merged['_merge'] == 'left_only'][['cities']]
# 筛选仅在df2存在的城市
only_in_df2 = merged[merged['_merge'] == 'right_only'][['cities']]

print("仅在df1存在的城市:")
print(only_in_df1)
print("仅在df2存在的城市:")
print(only_in_df2)

注意事项

如果cities列存在首尾空格(比如示例中的城市名),建议先清理数据:

df1['cities'] = df1['cities'].str.strip()
df2['cities'] = df2['cities'].str.strip()

内容的提问来源于stack exchange,提问作者Bruhlickd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 07:05:33