基于Python Pandas DataFrame两列值计算精确率与召回率
Hey there! Let's solve this problem efficiently—since you've got over 500 rows, we need a method that's both easy to implement and handles edge cases (like empty lists) without breaking.
Background: Standard Metric Definitions
First, let's align on the standard definitions to avoid confusion:
- Precision: The percentage of items in your actual output that are present in the expected output
Formula:(number of overlapping items) / (total items in actual output) - Recall: The percentage of items in your expected output that are correctly found in the actual output
Formula:(number of overlapping items) / (total items in expected output)
Step-by-Step Solution
Assuming your DataFrame is named df with columns Expected Output and Actual Output (each cell contains a list), here's how to compute the metrics in bulk:
1. Import Required Libraries
import pandas as pd
2. Define a Helper Function for Metrics
This function takes two lists, calculates their overlap, and returns precision/recall while avoiding division-by-zero errors:
def compute_precision_recall(expected_list, actual_list): # Convert lists to sets for fast intersection calculation expected_set = set(expected_list) actual_set = set(actual_list) # Count how many items appear in both lists overlap_count = len(expected_set.intersection(actual_set)) # Calculate precision (handle empty actual list to avoid errors) precision = overlap_count / len(actual_set) if len(actual_set) != 0 else 0.0 # Calculate recall (handle empty expected list to avoid errors) recall = overlap_count / len(expected_set) if len(expected_set) != 0 else 0.0 # Return results as a Series to easily add to the DataFrame return pd.Series([precision, recall], index=["Precision", "Recall"])
3. Apply the Function to Every Row
Use pandas.DataFrame.apply() to run the helper function on each row and add the results as new columns:
# Add Precision and Recall columns directly to your DataFrame df[["Precision", "Recall"]] = df.apply( lambda row: compute_precision_recall(row["Expected Output"], row["Actual Output"]), axis=1 )
Example Output
For your sample data, the resulting DataFrame will look like this:
| Expected Output | Actual Output | Precision | Recall |
|---|---|---|---|
| [2,10,5,266,8] | [7,2,9,266] | 0.5 | 0.4 |
| [4,89,34,453] | [4,22,34,453] | 0.75 | 0.75 |
Quick note on your initial example: You mentioned precision=2/5 for the first row, but that's not standard precision. Standard precision uses the actual output's length (4) as the denominator, giving 2/4=0.5. If you need a custom metric, just adjust the denominator in the function!
Why This Works
- Efficiency: Using sets for intersections is fast, even for 500+ rows.
- Robustness: Handles empty lists gracefully (returns 0 instead of throwing a division-by-zero error).
- Scalability: Works seamlessly as your dataset grows beyond 500 rows.
内容的提问来源于stack exchange,提问作者Mia

