Python 3读取特定格式txt多标签数据集的方法
Hey there, I've worked with similar multi-label dataset formats before – here's a clean Python solution to read your data and generate trainX and trainY exactly as you need:
Step-by-Step Code Implementation
def load_dataset(file_path): with open(file_path, 'r') as f: # Parse the header line to get dataset metadata header = f.readline().strip() total_rows, num_classes, dim = map(int, header.split()) trainX = [] trainY = [] # Process each data row for line_num, line in enumerate(f, start=2): # Start counting from 2 since header is line 1 line = line.strip() if not line: print(f"Skipping empty line {line_num}") continue # Split labels and features (split on first space) try: label_part, feature_part = line.split(' ', 1) except ValueError: print(f"Invalid line format at line {line_num}: {line}") continue # Convert labels to integer list for trainY labels = list(map(int, label_part.split(','))) trainY.append(labels) # Parse features into a dictionary, then generate ordered list for trainX feature_dict = {} for feat in feature_part.split(): idx_str, val_str = feat.split(':') try: idx = int(idx_str) val = float(val_str) feature_dict[idx] = val except ValueError: print(f"Invalid feature format at line {line_num}: {feat}") break else: # Ensure we have all dimensions (0 to dim-1) in order if len(feature_dict) != dim: print(f"Warning: Line {line_num} has {len(feature_dict)} features instead of {dim}") # Generate the feature list in correct order, fill missing values with 0 if needed x = [feature_dict.get(i, 0.0) for i in range(dim)] trainX.append(x) return trainX, trainY # Example usage trainX, trainY = load_dataset("your_dataset.txt") # Verify the first sample print("First sample features:", trainX[0]) print("First sample labels:", trainY[0])
Key Details Explained
- Header Parsing: We first read the first line to get the dataset's dimension count (120), which ensures we generate correctly sized feature lists.
- Label Handling: The comma-separated labels are converted directly into an integer list for
trainY– this preserves all multi-label entries exactly as they appear. - Feature Ordering: Since features are stored as "index:value" pairs (which might not be in order), we use a dictionary to map indices to values, then generate a list ordered from index 0 to 119. If any feature is missing, we default to
0.0(you can adjust this logic if missing values need different handling). - Error Handling: Added basic checks for invalid lines or feature formats to help debug any messy data entries.
Quick Validation
After running the code, trainX[0] will be a 120-element list of float values, and trainY[0] will be the list of labels (e.g., [11, 31, 65, 67]) from the first data row.
内容的提问来源于stack exchange,提问作者Ho Nam Cheung
相关产品推荐
相关产品推荐

