调用stats.percentileofscore处理多列时仅'b percent'列赋值成功,求助
问题原因与解决方案
核心问题
你的代码存在变量名冲突:外层定义了列表x = ['a','b'],但内层循环使用for x in x,第一次循环后x会被覆盖为字符串(依次变为'a'、'b')。后续行循环时,x已经是字符串而非列表,内层循环只会遍历字符串的单个字符,导致仅'b percent'列被持续赋值,'a percent'列仅在第一次循环时被处理。
修复方案
方案1:修正循环变量名
将外层列表变量名与内层循环变量名区分开,避免覆盖:
import pandas as pd from scipy import stats test = pd.DataFrame({'a':[20,2,7,1,7,7,34],'b':[100,99,102,103,56,70,200]}) test['a percent'] = '' test['b percent'] = '' # 重命名列表变量,避免冲突 cols = ['a','b'] for row in test.index: for col in cols: test.loc[row,f'{col} percent'] = stats.percentileofscore(test[col], test.loc[row,col])/100
方案2:更高效的Pandas风格写法
放弃嵌套循环,使用apply直接对列批量处理,代码更简洁且性能更好:
import pandas as pd from scipy import stats test = pd.DataFrame({'a':[20,2,7,1,7,7,34],'b':[100,99,102,103,56,70,200]}) for col in ['a','b']: # 对列中每个值计算百分位数,直接赋值给新列 test[f'{col} percent'] = test[col].apply(lambda val: stats.percentileofscore(test[col], val)/100)
修复后输出示例
a b a percent b percent 0 20 100 0.857143 0.571429 1 2 99 0.285714 0.428571 2 7 102 0.571429 0.714286 3 1 103 0.142857 0.857143 4 7 56 0.571429 0.142857 5 7 70 0.571429 0.285714 6 34 200 1.000000 1.000000
内容的提问来源于stack exchange,提问作者stvlam22
相关产品推荐
相关产品推荐

