You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决LightGBMError:标签754不小于标签映射数(31)?

LightGBM LambdaRank训练报错解决:Label 754 is not less than the number of label mappings (31)

问题重现

使用LightGBM训练LambdaMART(lambdarank目标)模型时,抛出错误:

LightGBMError: Label 754 is not less than the number of label mappings (31)

错误原因

  1. 标签范围不匹配:LightGBM的lambdarank目标默认num_class=31,但你的标签经过编码后最大值达到754,超出了默认的类别数量限制。
  2. 标签处理错误:代码中对包含relevance标签的整个DataFrame做了均值填充,把原本的整数标签转换成浮点数,破坏了标签的离散性。
  3. 变量混用:预测时使用了未定义的df_imputed,训练数据和评估数据的特征/标签处理逻辑不一致。

修复方案

1. 分离标签与特征的预处理

标签是离散的相关性等级,不需要做填充处理,仅对特征列进行缺失值填充:

# 只处理特征列,排除query_id和relevance
feature_cols = [col for col in df.columns if col not in ['query_id', 'relevance']]
df[feature_cols] = imputer.fit_transform(df[feature_cols])

2. 显式设置num_class参数

根据编码后的标签唯一数量,设置num_class参数让LightGBM知晓标签范围:

# 获取编码后标签的唯一类别数
num_classes = len(le.classes_)
params = {
    'objective': 'lambdarank',
    'metric': 'ndcg',
    'learning_rate': 0.05,
    'num_leaves': 31,
    'min_data_in_leaf': 20,
    'lambda_l1': 0.1,
    'lambda_l2': 0.1,
    'max_bin': 255,
    'num_iterations': 100,
    'ndcg_eval_at': [1, 3, 5],
    'num_class': num_classes  # 添加这一行
}

3. 修复变量混用问题

确保训练和评估使用同一批处理后的数据,替换未定义的df_imputed:

# 预测时使用处理后的result数据
predictions = ranker.predict(result.drop(['relevance', 'query_id'], axis=1))

# 评估时也使用result中的标签
ranking_accuracy = ndcg_score(
    [result['relevance'][result['query_id'] == q] for q in queries],
    [predictions[result['query_id'] == q] for q in queries]
)

4. 确保标签为整数类型

标签必须是整数,避免填充操作导致的类型转换问题:

# 确保标签是整数类型
result['relevance'] = result['relevance'].astype(int)

完整修正后的核心代码片段

import lightgbm as lgb
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import LabelEncoder
from sklearn.metrics import ndcg_score
import pandas as pd

def evaluate_population(population, file_path):
    if 'query_id' not in df.columns or 'relevance' not in df.columns:
        raise ValueError("DataFrame must contain 'query_id' and 'relevance' columns")
    
    evaluations = []
    imputer = SimpleImputer(strategy='mean')
    le = LabelEncoder()
    
    # 先编码标签
    df['relevance'] = le.fit_transform(df['relevance'])
    num_classes = len(le.classes_)
    
    # 仅对特征列做缺失值填充
    feature_cols = [col for col in df.columns if col not in ['query_id', 'relevance']]
    df[feature_cols] = imputer.fit_transform(df[feature_cols])
    
    for individual in population:
        selected_indices = [i for i, bit in enumerate(individual) if bit == '1']
        print(selected_indices)
        
        if not selected_indices:
            objective1 = len(file_path) + 1
            objective2 = 1.0
        else:
            queries = df['query_id'].unique()
            print(queries)
            selected_data = []

            for query in queries:
                query_data = df[df['query_id'] == query]
                num_docs = len(query_data)

                if len(individual) < num_docs:
                    individual = individual.ljust(num_docs, '0')

                mask = [int(bit) for bit in individual[:num_docs]]
                assert len(mask) == num_docs, "Mask length must match the number of documents"
                selected_docs = query_data.iloc[mask]
                selected_data.append(selected_docs)

            result = pd.concat(selected_data)
            # 确保标签是整数类型
            result['relevance'] = result['relevance'].astype(int)
            
            group = create_group(result['query_id'])
            train_data = lgb.Dataset(
                result.drop(['relevance', 'query_id'], axis=1), 
                label=result['relevance'], 
                group=group
            )

            params = {
                'objective': 'lambdarank',
                'metric': 'ndcg',
                'learning_rate': 0.05,
                'num_leaves': 31,
                'min_data_in_leaf': 20,
                'lambda_l1': 0.1,
                'lambda_l2': 0.1,
                'max_bin': 255,
                'num_iterations': 100,
                'ndcg_eval_at': [1, 3, 5],
                'num_class': num_classes
            }

            ranker = lgb.train(params, train_data)

            # 使用处理后的result数据预测
            predictions = ranker.predict(result.drop(['relevance', 'query_id'], axis=1))

            # 用result中的标签计算NDCG
            ranking_accuracy = ndcg_score(
                [result['relevance'][result['query_id'] == q] for q in queries],
                [predictions[result['query_id'] == q] for q in queries]
            )

            objective1 = len(selected_indices)
            objective2 = 1 - ranking_accuracy

        evaluations.append((individual, objective1, objective2))
    return evaluations

内容的提问来源于stack exchange,提问作者Amala K J

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 15:44:51