You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3读取特定格式txt多标签数据集的方法

Hey there, I've worked with similar multi-label dataset formats before – here's a clean Python solution to read your data and generate trainX and trainY exactly as you need:

Step-by-Step Code Implementation

def load_dataset(file_path):
    with open(file_path, 'r') as f:
        # Parse the header line to get dataset metadata
        header = f.readline().strip()
        total_rows, num_classes, dim = map(int, header.split())
        
        trainX = []
        trainY = []
        
        # Process each data row
        for line_num, line in enumerate(f, start=2):  # Start counting from 2 since header is line 1
            line = line.strip()
            if not line:
                print(f"Skipping empty line {line_num}")
                continue
            
            # Split labels and features (split on first space)
            try:
                label_part, feature_part = line.split(' ', 1)
            except ValueError:
                print(f"Invalid line format at line {line_num}: {line}")
                continue
            
            # Convert labels to integer list for trainY
            labels = list(map(int, label_part.split(',')))
            trainY.append(labels)
            
            # Parse features into a dictionary, then generate ordered list for trainX
            feature_dict = {}
            for feat in feature_part.split():
                idx_str, val_str = feat.split(':')
                try:
                    idx = int(idx_str)
                    val = float(val_str)
                    feature_dict[idx] = val
                except ValueError:
                    print(f"Invalid feature format at line {line_num}: {feat}")
                    break
            else:
                # Ensure we have all dimensions (0 to dim-1) in order
                if len(feature_dict) != dim:
                    print(f"Warning: Line {line_num} has {len(feature_dict)} features instead of {dim}")
                # Generate the feature list in correct order, fill missing values with 0 if needed
                x = [feature_dict.get(i, 0.0) for i in range(dim)]
                trainX.append(x)
        
        return trainX, trainY

# Example usage
trainX, trainY = load_dataset("your_dataset.txt")

# Verify the first sample
print("First sample features:", trainX[0])
print("First sample labels:", trainY[0])

Key Details Explained

  • Header Parsing: We first read the first line to get the dataset's dimension count (120), which ensures we generate correctly sized feature lists.
  • Label Handling: The comma-separated labels are converted directly into an integer list for trainY – this preserves all multi-label entries exactly as they appear.
  • Feature Ordering: Since features are stored as "index:value" pairs (which might not be in order), we use a dictionary to map indices to values, then generate a list ordered from index 0 to 119. If any feature is missing, we default to 0.0 (you can adjust this logic if missing values need different handling).
  • Error Handling: Added basic checks for invalid lines or feature formats to help debug any messy data entries.

Quick Validation

After running the code, trainX[0] will be a 120-element list of float values, and trainY[0] will be the list of labels (e.g., [11, 31, 65, 67]) from the first data row.

内容的提问来源于stack exchange,提问作者Ho Nam Cheung

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:31:40