Python DataFrame指定城市名清洗:避免非目标城市转为NaN
解决DataFrame城市名指定映射清洗问题
问题根源
你用map(cities_map)时,字典里没有的键会被转为NaN,这是Pandasmap方法的特性——仅保留字典中存在的映射,其余值替换为缺失值。要实现只修改指定城市别名,其余城市名不变,可以用以下两种简单方法:
方法1:使用replace方法(推荐)
replace方法专门针对值替换,只会处理字典中定义的键,未匹配的原数值保持不变。
# 定义目标城市及其别名 cities_names = ( ('Newton', ['west newton', 'newton center', 'chestnut hill', 'chestnuthill', 'waban', 'auberdale', 'auburndale']), ('Dover', ['dover ']), ('Needham', ['neeham']), ('Wellesley', ['wellesly']) ) # 构建替换字典:{别名: 目标城市名} cities_replace = {} for target_city, aliases in cities_names: for alias in aliases: cities_replace[alias] = target_city # 执行替换,仅修改指定别名,其余城市名保留原样 df_MA['City'] = df_MA['City'].replace(cities_replace) # 查看清洗后的唯一值 print(df_MA['City'].unique())
方法2:map结合fillna补全原数值
先用map处理指定别名,再把转为NaN的原数值填充回来:
cities_names = ( ('Newton', ['west newton', 'newton center', 'chestnut hill', 'chestnuthill', 'waban', 'auberdale', 'auburndale']), ('Dover', ['dover ']), ('Needham', ['neeham']), ('Wellesley', ['wellesly']) ) cities_map = {y: x[0] for x in cities_names for y in x[1]} # 先映射,再用原列值填充NaN df_MA['City'] = df_MA['City'].map(cities_map).fillna(df_MA['City'])
进阶:处理更多拼写错误(模糊匹配)
如果有大量不规则拼写错误,可使用fuzzywuzzy库做模糊匹配,自动识别相似的城市名:
from fuzzywuzzy import process def match_city(city_name, target_cities, threshold=80): # 匹配最相似的目标城市,相似度低于阈值则保留原名 best_match, score = process.extractOne(city_name, target_cities) return best_match if score >= threshold else city_name # 定义需要标准化的目标城市列表 target_cities = ['Newton', 'Dover', 'Needham', 'Wellesley'] # 应用模糊匹配清洗 df_MA['City'] = df_MA['City'].apply(lambda x: match_city(x, target_cities))
使用前需安装依赖:
pip install fuzzywuzzy python-Levenshtein
内容的提问来源于stack exchange,提问作者Orchid9
相关产品推荐
相关产品推荐

