You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python DataFrame指定城市名清洗:避免非目标城市转为NaN

解决DataFrame城市名指定映射清洗问题

问题根源

你用map(cities_map)时,字典里没有的键会被转为NaN,这是Pandasmap方法的特性——仅保留字典中存在的映射,其余值替换为缺失值。要实现只修改指定城市别名,其余城市名不变,可以用以下两种简单方法:


方法1:使用replace方法(推荐)

replace方法专门针对值替换,只会处理字典中定义的键,未匹配的原数值保持不变。

# 定义目标城市及其别名
cities_names = (
    ('Newton', ['west newton', 'newton center', 'chestnut hill', 'chestnuthill', 'waban', 'auberdale', 'auburndale']),
    ('Dover', ['dover ']), 
    ('Needham', ['neeham']), 
    ('Wellesley', ['wellesly'])
)

# 构建替换字典:{别名: 目标城市名}
cities_replace = {}
for target_city, aliases in cities_names:
    for alias in aliases:
        cities_replace[alias] = target_city

# 执行替换,仅修改指定别名,其余城市名保留原样
df_MA['City'] = df_MA['City'].replace(cities_replace)

# 查看清洗后的唯一值
print(df_MA['City'].unique())

方法2:map结合fillna补全原数值

先用map处理指定别名,再把转为NaN的原数值填充回来:

cities_names = (
    ('Newton', ['west newton', 'newton center', 'chestnut hill', 'chestnuthill', 'waban', 'auberdale', 'auburndale']),
    ('Dover', ['dover ']), 
    ('Needham', ['neeham']), 
    ('Wellesley', ['wellesly'])
)

cities_map = {y: x[0] for x in cities_names for y in x[1]}

# 先映射,再用原列值填充NaN
df_MA['City'] = df_MA['City'].map(cities_map).fillna(df_MA['City'])

进阶:处理更多拼写错误(模糊匹配)

如果有大量不规则拼写错误,可使用fuzzywuzzy库做模糊匹配,自动识别相似的城市名:

from fuzzywuzzy import process

def match_city(city_name, target_cities, threshold=80):
    # 匹配最相似的目标城市,相似度低于阈值则保留原名
    best_match, score = process.extractOne(city_name, target_cities)
    return best_match if score >= threshold else city_name

# 定义需要标准化的目标城市列表
target_cities = ['Newton', 'Dover', 'Needham', 'Wellesley']

# 应用模糊匹配清洗
df_MA['City'] = df_MA['City'].apply(lambda x: match_city(x, target_cities))

使用前需安装依赖:

pip install fuzzywuzzy python-Levenshtein

内容的提问来源于stack exchange,提问作者Orchid9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 05:33:30