You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何去除JSON中Word字段的重复项?现有代码无效求助

Fix Duplicate Word Entries in Speech Recognition JSON

Got it, let's tackle this duplicate Word issue in your speech recognition data. I've dealt with similar problems when cleaning up ASR outputs before, so here's a reliable approach that'll keep only the first occurrence of each Word and ditch the rest.

Core Idea

We'll use a set to track which words we've already encountered—sets let us check for existing entries super quickly. Then we'll loop through your original data list, adding each entry to a new list only if its Word hasn't been seen before.

Working Python Code

Let's assume your speech recognition JSON is parsed into a list of dictionaries (the most common format for this kind of data). Here's the code that'll do the trick:

import json

# Replace this with your actual speech recognition data
speech_recognition_results = [
    {"Word": "hi", "Confidence": 0.97},
    {"Word": "ask", "Confidence": 0.94},
    {"Word": "ask", "Confidence": 0.91},
    {"Word": "for", "Confidence": 0.99},
    {"Word": "help", "Confidence": 0.96},
    {"Word": "help", "Confidence": 0.93}
]

# Track words we've already kept
seen_words = set()
unique_results = []

for entry in speech_recognition_results:
    current_word = entry["Word"]
    # Only keep the entry if we haven't seen this word yet
    if current_word not in seen_words:
        seen_words.add(current_word)
        unique_results.append(entry)

# Print the cleaned data (or write it to a file)
print(json.dumps(unique_results, indent=2))

What This Does

  • seen_words acts as a "memory" of words we've already added to our cleaned list.
  • For each entry in your original data, we check if its Word is in seen_words. If not, we add the entry to unique_results and mark the word as seen.
  • This preserves the original order of entries and only keeps the first occurrence of each Word.

Example Output

Running the code above will give you this cleaned data:

[
  {
    "Word": "hi",
    "Confidence": 0.97
  },
  {
    "Word": "ask",
    "Confidence": 0.94
  },
  {
    "Word": "for",
    "Confidence": 0.99
  },
  {
    "Word": "help",
    "Confidence": 0.96
  }
]

If your previous code didn't work, it's likely because you were comparing entire dictionary entries instead of just the Word field, or using a method that doesn't preserve order (like converting to a set directly, which messes up sequence). This approach fixes both those issues.

内容的提问来源于stack exchange,提问作者Plumb InFront

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:27:20