如何基于两DataFrame的列匹配为CountryPoints添加Neighbour(Y/N)列
如何为DataFrame添加布尔列判断国家对是否为邻国
当然可行!你想要的其实是通过匹配两个DataFrame中的国家对,生成一个标记列,我给你分享两种简单高效的实现方法,都能得到你预期的结果。
首先我们先还原你的两个DataFrame:
import pandas as pd # 你的CountryPoints DataFrame country_points = pd.DataFrame({ 'From.country': ['Belgium', 'Belgium', 'Malta', 'Malta'], 'To.Country': ['Finland', 'Germany', 'Italy', 'UK'], 'points': [4, 5, 12, 1] }) # 相邻国家DataFrame neighbour_df = pd.DataFrame({ 'From.country': ['Belgium', 'Belgium', 'Malta'], 'To.Country': ['Finland', 'Germany', 'Italy'] })
方法一:左连接+填充替换
这种方法利用pandas的merge操作,先给邻国表标记Y,再左连接到主表,最后填充缺失值为N:
# 给邻国表添加标记列,然后左连接到主表 merged_result = country_points.merge( neighbour_df.assign(Neighbour='Y'), on=['From.country', 'To.Country'], how='left' ) # 将未匹配到的NaN替换为'N' merged_result['Neighbour'] = merged_result['Neighbour'].fillna('N')
左连接会保留主表的所有行,匹配到邻国表的行会带上Y,没匹配到的则是NaN,最后用fillna把NaN换成N就完成了。
方法二:利用元组集合匹配
这种方法更直观,先把邻国的国家对转换成元组集合,然后判断主表每行的国家对是否在集合中:
# 将邻国表的国家对转换成元组集合,方便快速查找 neighbour_pairs = set(zip(neighbour_df['From.country'], neighbour_df['To.Country'])) # 生成Neighbour列:匹配到就返回Y,否则返回N country_points['Neighbour'] = ( country_points[['From.country', 'To.Country']] .apply(tuple, axis=1) .isin(neighbour_pairs) .map({True: 'Y', False: 'N'}) )
zip把两列转换成元组对,set让查找更高效;apply(tuple, axis=1)把主表每行的两个国家也转成元组,isin判断是否在集合中,最后用map把布尔值转换成Y/N。
两种方法最终都会得到你想要的结果:
From.country To.Country points Neighbour 0 Belgium Finland 4 Y 1 Belgium Germany 5 Y 2 Malta Italy 12 Y 3 Malta UK 1 N
内容的提问来源于stack exchange,提问作者user9737581
相关产品推荐
相关产品推荐

