如何用Python计算已量化为10级的数据集各特征列的熵?
Hey there! Let's fix up your entropy calculation code step by step. First, let's recap your setup: you've got a dataset quantized into 10 levels (0-9), sample data like 9 9 1 8 9 1 1 9 3 6 1 0 8 3 8 4 4 1 0 2 1 9 9 0, where the first 5 values belong to category 1, and you need to compute entropy for each feature column.
Issues in Your Original Code
Let's break down the problems first:
- Redundant file handling: You open the file with
open()then pass it topd.read_csv()—pandas can handle file paths directly, so this extra step is unnecessary. - Incomplete logic: The
df.loc[:,...line cuts off, so we're missing the core code to count value frequencies and compute entropy. - Separator mismatch: Your sample data uses spaces as separators, but your code specifies
sep='\t'(tabs), which would load the data incorrectly.
Corrected Code with Explanations
Here's a complete, working version tailored to your needs:
import pandas as pd import math # Load the dataset directly (pandas handles file paths natively) # Use sep='\s+' to handle any number of spaces as separators # Adjust the column count in names if your dataset has more/less than 8 features df = pd.read_csv('data1.txt', sep='\s+', header=None, names=[f'feature_{i+1}' for i in range(8)]) def calculate_entropy(column): """Calculate Shannon Entropy for a single feature column""" # Get frequency of each quantized level in the column value_counts = column.value_counts() total_samples = len(column) entropy = 0.0 for count in value_counts: # Calculate probability of the current level prob = count / total_samples # Avoid log(0) errors (only run if probability is positive) if prob > 0: entropy -= prob * math.log2(prob) return entropy # Compute entropy for every feature column entropy_results = {} for col_name in df.columns: entropy_results[col_name] = calculate_entropy(df[col_name]) # Print readable results print("Entropy for each feature column:") for col, entropy_val in entropy_results.items(): print(f"{col}: {entropy_val:.4f}")
Key Fixes & Notes
- File Loading: Swapped
sep='\t'forsep='\s+'to match your space-separated sample data. The column names are generated dynamically to avoid manual typing. - Entropy Function: Implements the standard Shannon Entropy formula ( H = -\sum_{i=1}^{n} p_i \log_2(p_i) ):
- Uses
value_counts()to count how often each 0-9 level appears in a column - Calculates probability for each level
- Adds a guard clause to skip log(0) (which would throw an error)
- Uses
- Column Iteration: Loops through every feature column, computes its entropy, and stores results in a dictionary for easy reference.
Testing with Your Sample Data
If your data1.txt splits the sample 24 values into 3 rows of 8 features each:
9 9 1 8 9 1 1 9 3 6 1 0 8 3 8 4 4 1 0 2 1 9 9 0
Running the code will output entropy values like:
Entropy for each feature column: feature_1: 1.5850 feature_2: 1.5850 feature_3: 0.9183 ...
Bonus: Conditional Entropy (If Needed)
You mentioned the first 5 values belong to category 1—if you need to compute entropy of features given category labels (conditional entropy), just let me know and we can adjust the code for that!
内容的提问来源于stack exchange,提问作者Amir

