如何批量替换DataFrame中城市的非数字开头Code为对应数字编码?
解决方案
针对你的需求,这里有一个高效且无需逐个城市编写代码的方法,适合处理大规模数据:
步骤1:创建测试数据(和你提供的一致)
import pandas as pd import numpy as np df = {'City': ['London','Tokyo','London','Paris','Paris','London','Tokyo','Tokyo', 'Paris','Berlin','Berlin','Berlin'], 'Code': ['367','812','367','964','964','BN611','812','Y366','Z167','L715','412','L715']} df = pd.DataFrame(data=df)
步骤2:构建城市与纯数字编码的映射字典
先筛选出每个城市对应的唯一纯数字编码,构建映射关系:
# 筛选纯数字编码的行,去重后生成城市-编码的字典 city_code_mapping = df[df['Code'].str.isdigit()].drop_duplicates('City').set_index('City')['Code'].to_dict()
步骤3:批量替换非纯数字编码
用矢量化操作快速完成替换,效率远高于逐行遍历:
# 判断Code是否为纯数字,非纯数字则用对应城市的基准编码替换 df['Code'] = np.where(df['Code'].str.isdigit(), df['Code'], df['City'].map(city_code_mapping))
最终处理结果
| City | Code |
|---|---|
| London | 367 |
| Tokyo | 812 |
| London | 367 |
| Paris | 964 |
| Paris | 964 |
| London | 367 |
| Tokyo | 812 |
| Tokyo | 812 |
| Paris | 964 |
| Berlin | 412 |
| Berlin | 412 |
| Berlin | 412 |
关键说明
str.isdigit()精准匹配你的数据类型(仅纯数字、字母开头两类编码)- 映射字典自动提取每个城市的纯数字编码,无需手动指定
np.where是矢量化操作,适合处理数千甚至更多行的大数据量,性能优于apply逐行处理
内容的提问来源于stack exchange,提问作者Fold_In_The_Cheese
相关产品推荐
相关产品推荐

