Pandas重采样触发ValueError: a must be greater than 0问题求助
问题解决:重采样触发
ValueError: a must be greater than 0 unless no samples are taken 问题背景
尝试对重平衡数据集churn_train中Churn列值为True的样本重采样,抽取158条记录,但触发报错。已确认数据集非空,Churn为True的样本共320条。
相关信息
数据集片段(churn部分行)
State,Account Length,Area Code,Phone,Intl Plan,VMail Plan,VMail Message,Day Mins,Day Calls,Day Charge,Eve Mins,Eve Calls,Eve Charge,Night Mins,Night Calls,Night Charge,Intl Mins,Intl Calls,Intl Charge,CustServ Calls,Old Churn,Churn "KS",128,415,"382-4657","no","yes",25,265.100000,110,45.070000,197.400000,99,16.780000,244.700000,91,11.010000,10.000000,3,2.700000,1,"False.","False" "OH",107,415,"371-7191","no","yes",26,161.600000,123,27.470000,195.500000,103,16.620000,254.400000,103,11.450000,13.700000,3,3.700000,1,"False.","False" "IN",65,415,"329-6603","no","no",0,129.100000,137,21.950000,228.500000,83,19.420000,208.800000,111,9.400000,12.700000,6,3.430000,4,"True.","True"
执行代码
churn_train['Churn'].value_counts() # 输出: # False 1913 # True 320 # Name: Churn, dtype: int64 to_resample = churn_train.loc[churn_train['Churn'] == "True"] our_resample = to_resample.sample(n = 158, replace = True) churn_train_rebal = pd.concat([churn_train, our_resample])
报错信息
ValueError Traceback (most recent call last) /var/folders/wv/42dn23fd1cb0czpvqdnb6zw00000gn/T/ipykernel_7751/2929105044.py in <module> 1 to_resample = churn_train.loc[churn_train['Churn'] == "True"] ----> 2 our_resample = to_resample.sample(n = 158, replace = True) 3 churn_train_rebal = pd.concat([churn_train, our_resample]) ~/opt/miniconda3/lib/python3.9/site-packages/pandas/core/generic.py in sample(self, n, frac, replace, weights, random_state, axis, ignore_index) 5452 weights = sample.preprocess_weights(self, weights, axis) 5453 -> 5454 sampled_indices = sample.sample(obj_len, size, replace, weights, rs) 5455 result = self.take(sampled_indices, axis=axis) 5456 ~/opt/miniconda3/lib/python3.9/site-packages/pandas/core/sample.py in sample(obj_len, size, replace, weights, random_state) 148 raise ValueError("Invalid weights: weights sum to zero") 149 -> 150 return random_state.choice(obj_len, size=size, replace=replace, p=weights).astype( 151 np.intp, copy=False 152 ) mtrand.pyx in numpy.random.mtrand.RandomState.choice() ValueError: a must be greater than 0 unless no samples are taken
问题原因
从value_counts的输出可以看到,Churn列的取值是布尔类型的True/False,但代码中用了字符串"True"进行匹配,导致筛选出的to_resample是空DataFrame。空DataFrame调用sample(n=158)时,就会触发ValueError——因为没有可采样的样本。
解决方法
将筛选条件修改为匹配布尔值True,而非字符串"True":
# 方法1:直接匹配布尔值 to_resample = churn_train.loc[churn_train['Churn'] == True] # 方法2:更简洁的布尔索引写法 to_resample = churn_train[churn_train['Churn']] our_resample = to_resample.sample(n = 158, replace = True) churn_train_rebal = pd.concat([churn_train, our_resample])
如果不确定列的数据类型,可以先执行print(churn_train['Churn'].dtype)确认类型,再调整匹配条件。
内容的提问来源于stack exchange,提问作者c200402
相关产品推荐
相关产品推荐

