如何在np.where中对DataFrame单行字段提取数字首字符?
问题:在Pandas中使用np.where实现按行条件生成新列
先创建如下Pandas DataFrame:
import pandas as pd import numpy as np data = [10,20,30,40,50,60] df = pd.DataFrame(data, columns=['Numbers']) df['add_7'] = (df['Numbers'] + 7)
生成的DataFrame如下:
| Numbers | add_7 |
|---|---|
| 10 | 17 |
| 20 | 27 |
| 30 | 37 |
| 40 | 47 |
| 50 | 57 |
| 60 | 67 |
需求是创建名为first_digit的新列:
- 当
add_7列的值是3的倍数时,取Numbers列对应行数字的首字符(转为字符串) - 否则填入"not a multiple of three"
期望结果:
| Numbers | add_7 | first_digit |
|---|---|---|
| 10 | 17 | not a multiple of three |
| 20 | 27 | 2 |
| 30 | 37 | not a multiple of three |
| 40 | 47 | not a multiple of three |
| 50 | 57 | 5 |
| 60 | 67 | not a multiple of three |
尝试了以下代码,但结果错误:
df['first_digit'] = np.where(df['add_7'] % 3 == 0, str(df['Numbers'][0]), 'not a multiple of three')
错误结果:
| Numbers | add_7 | first_digit |
|---|---|---|
| 10 | 17 | not a multiple of three |
| 20 | 27 | 10 |
| 30 | 37 | not a multiple of three |
| 40 | 47 | not a multiple of three |
| 50 | 57 | 10 |
| 60 | 67 | not a multiple of three |
问题:如何在np.where中指定仅对当前行的字段进行操作,而非整个列?
解决方案
错误原因
你之前的代码里str(df['Numbers'][0])是取Numbers列第一行的值转为字符串,然后把这个固定值套用在所有满足条件的行,才会出现所有符合条件的行都显示10的错误。要实现按行操作,需要对整个Numbers列做批量的首字符提取,而非取单个值。
正确代码
先对Numbers列批量提取首字符生成Series,再传入np.where中:
# 批量生成每行Numbers的首字符Series first_chars = df['Numbers'].astype(str).str[0] # 按条件赋值 df['first_digit'] = np.where(df['add_7'] % 3 == 0, first_chars, 'not a multiple of three')
也可以合并为一行代码:
df['first_digit'] = np.where( df['add_7'] % 3 == 0, df['Numbers'].astype(str).str[0], 'not a multiple of three' )
代码说明
df['Numbers'].astype(str).str[0]:将Numbers列的每个数值转为字符串,再提取第一个字符,这一步是按行批量处理,得到的Series和原DataFrame行数一致,每个元素对应原行的首字符。np.where接收这个Series后,会自动匹配满足条件的行,填入对应位置的首字符,而非固定值。
运行结果
执行后得到的DataFrame与期望结果一致:
| Numbers | add_7 | first_digit |
|---|---|---|
| 10 | 17 | not a multiple of three |
| 20 | 27 | 2 |
| 30 | 37 | not a multiple of three |
| 40 | 47 | not a multiple of three |
| 50 | 57 | 5 |
| 60 | 67 | not a multiple of three |
内容的提问来源于stack exchange,提问作者kd8
相关产品推荐
相关产品推荐

