pandas Series执行字符串函数报错 筛选数据提示非字符串数组错误
DataFrame筛选字符串开头记录报错解决
问题复现代码
构造测试数据集:
import pandas as pd import numpy as np students = [('jack', 34, 'Sydeny', 'Australia'), ('Riti', 30, 'Delhi', 'India'), ('Vikas', 31, 'Mumbai', 'India'), ('Neelu', 32, 'Bangalore', 'India'), ('John', 16, 'New York', 'US'), ('Mike', 17, 'las vegas', 'US')] df = pd.DataFrame( students, columns=['Name', 'Age', 'City', 'Country'], index=['a', 'b', 'c', 'd', 'e', 'f'])
尝试筛选Country字段以'I'开头的记录的代码:
print(df.loc[lambda x:np.char.startswith(x['Country'],'I')])
运行抛出错误:
string operation on non-string array
尝试修复的无效操作:
df.astype({'Country':str})
执行后错误依旧存在。
出错原因
np.char.startswith仅支持numpy原生字符串类型(np.str_/固定长度字符串dtype)的数组。pandas默认存储字符串的列是objectdtype,底层是Python原生字符串对象组成的数组,不属于numpy字符串数组,直接传入就会触发非字符串数组报错。df.astype({'Country':str})不会原地修改DataFrame,pandas默认返回类型转换后的新对象,没有赋值给原变量的话转换操作完全不生效;就算赋值回原变量,转成Pythonstr类型后列的dtype依然是object,还是无法被np.char.startswith识别。
正确实现方式
优先使用pandas原生字符串访问器实现,不需要额外做类型转换,是最稳定的写法:
# 写法1:loc搭配lambda,适合链式调用场景 print(df.loc[lambda x: x['Country'].str.startswith('I')]) # 写法2:直接布尔索引,写法更简洁 print(df[df['Country'].str.startswith('I')])
如果一定要使用np.char.startswith,需要先将列转为numpy字符串数组再传入,同时注意类型转换要赋值生效:
# 先将列转为numpy支持的字符串类型,再提取numpy数组传入 df['Country'] = df['Country'].astype(np.str_) print(df.loc[lambda x: np.char.startswith(x['Country'].to_numpy(), 'I')])
以上代码运行后都会正确返回所有Country字段以'I'开头的3条记录。
内容的提问来源于stack exchange,提问作者user2779311
相关产品推荐
相关产品推荐

