使用PuLP求解最大化正确分类的线性优化问题:阈值异常导致无正类分类的问题求助
看起来你遇到的核心问题是约束条件的逻辑搞反了,导致模型为了满足“最多100个误分类0”的限制,直接选择了最保守的可行解——把阈值设得极高,完全不触发正类分类。咱们一步步拆解问题,再修正代码:
一、问题出在这几个关键约束上
1. Z变量的逻辑约束完全写反了
你定义Z[j,i]是“样本j的特征i是否超过阈值”,但原来的约束Z[j, i] >= (data.iloc[j, i] - T[i])/M_large方向完全错误。这个约束会要求当data[j,i] >= T[i]时,Z必须大于等于一个正数,但Z是0/1变量,很多情况下无法满足,模型只能反过来把T[i]设得远大于data[j,i],让data-T变成负数,这样约束自动满足(Z>=负数,而Z本身非负),结果就是Z全为0,自然没有样本被分类为1。
正确的做法是用大M法来表达“如果特征i被选中,那么Z[j,i]等于1当且仅当data[j,i] >= T[i]”:
- 当X[i]=1时:
- 若data[j,i] >= T[i],则Z[j,i]必须为1 →
T[i] ≤ data.iloc[j,i] + M_large*(1 - Z[j,i]) - 若data[j,i] < T[i],则Z[j,i]必须为0 →
data.iloc[j,i] ≤ T[i] + M_large*Z[j,i]
- 若data[j,i] >= T[i],则Z[j,i]必须为1 →
- 当X[i]=0时,Z[j,i]必须为0 →
Z[j,i] ≤ X[i](这部分你原来的是对的)
这里的M_large要选一个比数据范围大的数,比如你的数据是0-1之间的随机数,设成2就足够安全。
2. 阈值的上下界设置太宽松
你给T[i]设了-100到100的范围,但你的数据全是0-1,模型完全可以把阈值设到1以上,这样没有任何样本能超过阈值,自然不会有Y[j]=1的情况。把T[i]的上下界限制在数据的实际范围内(比如0到1),能避免这种极端情况。
3. 约束未考虑“未选中特征不参与逻辑”
虽然你有Z[j,i] ≤ X[i],但最好在T的约束里也加入(1-X[i])*M_large的项,确保未选中的特征不会影响Z变量的逻辑,让模型更清晰。
二、修改后的完整代码
import pandas as pd import numpy as np from pulp import LpMaximize, LpProblem, LpVariable, lpSum, value # Sample data np.random.seed(42) data = pd.DataFrame(np.random.rand(1000, 10), columns=[f'X{i}' for i in range(10)]) # 10 continuous indicators data['target'] = np.random.choice([0, 1], size=1000, p=[0.7, 0.3]) # Binary target (30% ones) # Parameters M = 5 # Maximum number of selected indicators S = 3 # Minimum number of indicators above threshold to classify as 1 max_false_positives = 100 # Maximum allowed misclassified 0s # Get data bounds for T variables data_min = data.min() data_max = data.max() M_large = 2 # Larger than the max range of data (0-1) # Define model model = LpProblem("Indicator_Selection", LpMaximize) # Decision variables X = {i: LpVariable(f'X_{i}', cat='Binary') for i in range(10)} # Indicator selection # Restrict T to data's actual range to avoid extreme values T = {i: LpVariable(f'T_{i}', lowBound=data_min[f'X{i}'], upBound=data_max[f'X{i}'], cat='Continuous') for i in range(10)} Y = {j: LpVariable(f'Y_{j}', cat='Binary') for j in range(len(data))} # Predicted labels # Auxiliary binary variables to represent whether each indicator exceeds its threshold Z = {(j, i): LpVariable(f'Z_{j}_{i}', cat='Binary', lowBound=0, upBound=1) for j in range(len(data)) for i in X} # Constraint: Select at most M indicators model += lpSum(X[i] for i in X) <= M # Constraints for Z[j,i]: link to data, T, and X[i] for j in range(len(data)): for i in X: # If X[i] is 1, Z[j,i] = 1 iff data[j,i] >= T[i] # 1. Ensure T[i] <= data[j,i] when Z[j,i] = 1 (and X[i] =1) model += T[i] <= data.iloc[j, i] + M_large * (1 - Z[j,i]) + M_large * (1 - X[i]) # 2. Ensure data[j,i] <= T[i] when Z[j,i] =0 (and X[i] =1) model += data.iloc[j, i] <= T[i] + M_large * Z[j,i] + M_large * (1 - X[i]) # 3. If X[i] is 0, Z[j,i] must be 0 model += Z[j,i] <= X[i] # Constraint: Classification rule (Y[j] =1 only if at least S indicators are above threshold) for j in range(len(data)): model += lpSum(Z[j,i] for i in X) >= S * Y[j] # Constraint: Limit false positives (misclassified 0s) false_positives = lpSum(Y[j] for j in range(len(data)) if data.iloc[j, 'target'] == 0) model += false_positives <= max_false_positives # Objective: Maximize true positives (correctly classified 1s) true_positives = lpSum(Y[j] for j in range(len(data)) if data.iloc[j, 'target'] == 1) model += true_positives # Solve the model model.solve() # Output results selected_features = [f'X{i}' for i in X if value(X[i]) == 1] thresholds = {f'X{i}': round(value(T[i]), 4) for i in X if value(X[i]) == 1} true_positives_count = int(value(true_positives)) false_positives_count = int(value(false_positives)) print(f"Selected features: {selected_features}") print(f"Thresholds: {thresholds}") print(f"True positives (correct 1s): {true_positives_count}") print(f"False positives (misclassified 0s): {false_positives_count}") # Check first 10 sample predictions sample_predictions = [int(value(Y[j])) for j in range(10)] print(f"First 10 predictions: {sample_predictions}")
三、修改后能解决问题的原因
- Z变量的约束逻辑纠正:现在模型能正确判断样本是否超过阈值,不会再被迫设置极端阈值。
- 阈值范围限制:T[i]被约束在数据的0-1范围内,避免了模型选到1以上的阈值。
- 未选中特征的隔离:通过
M_large*(1-X[i])项,确保未选中的特征不会干扰Z变量的逻辑判断。
运行这段代码后,你会看到模型会选中M个特征,设置合理的阈值,并且会有一定数量的正类被正确分类,同时误分类的0数量也会控制在100以内。
四、额外提示
因为你的样本数据是随机生成的,特征和target没有实际相关性,所以模型的分类效果可能不会特别好,但至少不会出现“没有样本被分类为1”的情况。如果是真实数据,特征和target有相关性,效果会更明显。
备注:内容来源于stack exchange,提问作者Marco Ballerini

