基于Pandas DataFrame某列首字符填充新列的实现方法
实现Pandas按type首字符映射添加category列的方法
嘿,我来帮你搞定这个Pandas的映射需求!先理清楚你的场景:你有这样一个原始DataFrame:
+-----+------+ | No | type | +-----+------+ | 123 | C01 | | 123 | C02 | | 123 | T01 | | 345 | C01 | | 345 | H12 | | 345 | H22 | +-----+------+
还有一个NumPy数组 arr = ["Car", "Tree", "House"],需要根据type列的首字符(C对应Car,T对应Tree,H对应House)添加category列,得到目标结果。下面给你几种简单好用的实现方法:
方法一:字典映射 + 字符串提取首字符
这是最直观的方法,先创建一个映射字典把首字符和对应类别绑定,再提取type的首字符做映射:
import pandas as pd import numpy as np # 构造你的原始DataFrame df = pd.DataFrame({ 'No': [123, 123, 123, 345, 345, 345], 'type': ['C01', 'C02', 'T01', 'C01', 'H12', 'H22'] }) arr = ["Car", "Tree", "House"] # 建立首字符到类别的映射字典 mapping_dict = {'C': arr[0], 'T': arr[1], 'H': arr[2]} # 提取type列的首字符,并用map完成映射,添加新列 df['category'] = df['type'].str[0].map(mapping_dict) # 查看结果 print(df)
运行后就能得到你想要的输出:
No type category 0 123 C01 Car 1 123 C02 Car 2 123 T01 Tree 3 345 C01 Car 4 345 H12 House 5 345 H22 House
方法二:正则替换法
如果你不想写字典,也可以用str.replace()结合正则表达式,直接匹配开头字符并替换:
import pandas as pd import numpy as np df = pd.DataFrame({ 'No': [123, 123, 123, 345, 345, 345], 'type': ['C01', 'C02', 'T01', 'C01', 'H12', 'H22'] }) arr = ["Car", "Tree", "House"] # 依次替换以C、T、H开头的type值为对应类别 df['category'] = df['type'].str.replace(r'^C.*', arr[0]) df['category'] = df['category'].str.replace(r'^T.*', arr[1]) df['category'] = df['category'].str.replace(r'^H.*', arr[2])
这个方法的思路是:匹配以C/T/H开头的整个字符串,直接替换成对应的类别,操作起来也很简单。
方法三:多条件映射(适合扩展)
如果后续你有更多的映射规则,用np.select()会更灵活,能轻松处理多条件场景:
import pandas as pd import numpy as np df = pd.DataFrame({ 'No': [123, 123, 123, 345, 345, 345], 'type': ['C01', 'C02', 'T01', 'C01', 'H12', 'H22'] }) arr = ["Car", "Tree", "House"] # 定义匹配条件:分别匹配以C、T、H开头的type值 conditions = [ df['type'].str.startswith('C'), df['type'].str.startswith('T'), df['type'].str.startswith('H') ] # 对应条件的结果选项 choices = arr # 根据条件匹配结果添加category列 df['category'] = np.select(conditions, choices)
这种方法把条件和结果分开定义,后续新增规则只需要在conditions和choices里加内容就行,维护起来很方便。
内容的提问来源于stack exchange,提问作者MaMo
相关产品推荐
相关产品推荐

