You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为DataFrame列追加列索引字符串及Learning to Rank数据集咨询

Adding Column Index Strings to Your Letor-Style LTR Dataset in Pandas

Hey there! I’ve worked with Letor datasets for Learning to Rank projects before, so I can walk you through how to properly assign meaningful column names to your pandas DataFrame. Let’s break this down step by step:

Step 1: Load the Raw Dataset

First, we’ll read the space-separated text file into a DataFrame. Since the dataset uses one or more spaces as separators, we’ll use sep='\s+' in pd.read_csv:

import pandas as pd

# Replace 'your_ltr_data.txt' with your actual file path
df = pd.read_csv('your_ltr_data.txt', sep='\s+', header=None)

Step 2: Clean Up the Query ID Column

The second column in your data has the format qid:X — we’ll extract just the numeric part to make it usable:

# Extract the numeric query ID from the 'qid:X' string
df[1] = df[1].str.extract('qid:(\d+)').astype(int)

Step 3: Generate and Assign Column Names

We’ll create a list of column names tailored to your dataset structure:

  • First column: rank (your initial ranking value)
  • Second column: query_id (the cleaned numeric query ID)
  • Remaining columns: feature_1, feature_2, ..., feature_46 (matching your 46 features)

Option 1: Fixed Feature Count (46 features)

If you know you always have 46 features, use this straightforward approach:

# Create column names list
column_names = ['rank', 'query_id'] + [f'feature_{i}' for i in range(1, 47)]

# Assign names to the DataFrame
df.columns = column_names

Option 2: Dynamic Feature Count (for variable feature numbers)

If your dataset might have a different number of features, calculate it dynamically based on the DataFrame’s shape:

# Calculate number of features (total columns minus 2 for rank and query_id)
num_features = df.shape[1] - 2

# Generate column names
column_names = ['rank', 'query_id'] + [f'feature_{i}' for i in range(1, num_features + 1)]

# Assign names
df.columns = column_names

Step 4: Verify the Result

Check the first few rows to make sure everything looks right:

print(df.head())

Quick Tip

If you run into issues with malformed lines in the dataset, add on_bad_lines='skip' to pd.read_csv to skip problematic rows (just make sure you’re okay losing that data!).

内容的提问来源于stack exchange,提问作者Darren Christopher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 07:04:12