Pandas数据框中对比Last_Price与Price列末位数字并设置Marked列的问题
解决方案
问题原因
你的Last_Price和Price列是浮点型(float),直接调用字符串方法strip()会触发AttributeError;而str(subdf['Price'])是把整个Series转为字符串,无法获取单个元素的末位数字。
正确实现代码
使用Pandas的矢量化字符串操作(避免低效循环),先将列转为字符串,再提取每个元素的最后一位进行对比:
import pandas as pd import numpy as np # 提取两列每个值的最后一位数字(转为字符串后取末尾) last_price_end = subdf['Last_Price'].astype(str).str[-1] price_end = subdf['Price'].astype(str).str[-1] # 对比末位:不相等则赋值'X',否则设为NaN subdf['Marked'] = np.where(last_price_end != price_end, 'X', np.nan)
处理特殊情况(可选)
如果数据中存在类似2.80这类末尾带0的数值,转字符串会得到"2.80",末位是0。若你需要忽略末尾的0和小数点(取有效数字的最后一位),可以用apply配合格式化处理:
def get_last_digit(num): # 格式化后去掉末尾的0和小数点,再取最后一位 cleaned = f"{num:.10f}".rstrip('0').rstrip('.') return cleaned[-1] last_price_end = subdf['Last_Price'].apply(get_last_digit) price_end = subdf['Price'].apply(get_last_digit) subdf['Marked'] = np.where(last_price_end != price_end, 'X', np.nan)
测试结果
针对你的示例数据,运行后Marked列结果如下:
| Last_Price | Price | Marked |
|---|---|---|
| 3 | 2.89 | X |
| 1.99 | 2.09 | NaN |
| 3.9 | 3.79 | X |
内容的提问来源于stack exchange,提问作者Dritan
相关产品推荐
相关产品推荐

