Python新手求助:列表去重pred值,保留第7位浮点值最高项
Hey there! It totally makes sense to feel stuck when you're new to Python and dealing with data deduplication based on specific criteria—let's break this down step by step.
First, let's clarify your goal: you have a list where each item has 7 values, you want to remove duplicates based on the first value (pred), and for any duplicates, keep only the item with the highest value in the 7th position (a float). Your current list comprehension generates items with 4 values, so I'll adjust the solution to work with both scenarios (just tweak the score index if needed!).
Step-by-Step Solution
The core idea is to use a dictionary to group items by pred—since dictionary keys are unique, this automatically handles deduplication. We'll update the dictionary entry for each pred whenever we find an item with a higher target float value.
- Generate your initial list using your existing code (or use your original 7-element list if that's your source data):
# Your original list comprehension (produces 4-element tuples) temp_list = [(pred, loc, rec, score) if rec else ("unknown", loc, False, 0) for pred, loc, rec, score in zip(knn_clf.predict(faces_encodings), X_face_locations, are_matches, scores)]
- Use a dictionary to track the highest-score item per
pred:
# Initialize an empty dictionary to map each pred to its best (highest-score) item pred_best_map = {} for item in temp_list: current_pred = item[0] # If your items have 7 values, change this index from 3 to 6 (Python uses 0-based indexing) current_score = item[3] # Add the item to the dictionary if it's the first occurrence of the pred, # or replace the existing entry if the current item has a higher score if current_pred not in pred_best_map or current_score > pred_best_map[current_pred][3]: pred_best_map[current_pred] = item
- Convert the dictionary back to a list:
# This gives you the final deduplicated list with only the highest-score item per pred final_result = list(pred_best_map.values()) return final_result
How This Works
- The dictionary
pred_best_mapensures eachpredonly appears once, eliminating duplicates by design. - For every item, we either add it to the dictionary (if it's the first time we see that
pred) or replace the existing entry if the current item has a higher score. - Converting the dictionary's values to a list gives you your cleaned-up, deduplicated data.
If your original items are indeed 7-element tuples, just adjust the score index from 3 to 6 in the code—that's the only change you need!
内容的提问来源于stack exchange,提问作者Michael Mano

