如何去除JSON中Word字段的重复项?现有代码无效求助
Word Entries in Speech Recognition JSON Got it, let's tackle this duplicate Word issue in your speech recognition data. I've dealt with similar problems when cleaning up ASR outputs before, so here's a reliable approach that'll keep only the first occurrence of each Word and ditch the rest.
Core Idea
We'll use a set to track which words we've already encountered—sets let us check for existing entries super quickly. Then we'll loop through your original data list, adding each entry to a new list only if its Word hasn't been seen before.
Working Python Code
Let's assume your speech recognition JSON is parsed into a list of dictionaries (the most common format for this kind of data). Here's the code that'll do the trick:
import json # Replace this with your actual speech recognition data speech_recognition_results = [ {"Word": "hi", "Confidence": 0.97}, {"Word": "ask", "Confidence": 0.94}, {"Word": "ask", "Confidence": 0.91}, {"Word": "for", "Confidence": 0.99}, {"Word": "help", "Confidence": 0.96}, {"Word": "help", "Confidence": 0.93} ] # Track words we've already kept seen_words = set() unique_results = [] for entry in speech_recognition_results: current_word = entry["Word"] # Only keep the entry if we haven't seen this word yet if current_word not in seen_words: seen_words.add(current_word) unique_results.append(entry) # Print the cleaned data (or write it to a file) print(json.dumps(unique_results, indent=2))
What This Does
seen_wordsacts as a "memory" of words we've already added to our cleaned list.- For each entry in your original data, we check if its
Wordis inseen_words. If not, we add the entry tounique_resultsand mark the word as seen. - This preserves the original order of entries and only keeps the first occurrence of each
Word.
Example Output
Running the code above will give you this cleaned data:
[ { "Word": "hi", "Confidence": 0.97 }, { "Word": "ask", "Confidence": 0.94 }, { "Word": "for", "Confidence": 0.99 }, { "Word": "help", "Confidence": 0.96 } ]
If your previous code didn't work, it's likely because you were comparing entire dictionary entries instead of just the Word field, or using a method that doesn't preserve order (like converting to a set directly, which messes up sequence). This approach fixes both those issues.
内容的提问来源于stack exchange,提问作者Plumb InFront

