You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对字符串中的特定词汇进行扩展?附待扩展词汇列表

How to Expand or Normalize Specific Vocabulary in a String

Great question! Let's walk through how to work with your synonym lists to either normalize (standardize) or expand terms in a string. I'll use Python for examples since it's widely used for text processing, but the logic translates to other languages too.

First, let's restructure your separate lists into a more manageable dictionary—this will make it easier to map synonyms to their target terms or groups:

# Organize your synonym groups with a canonical (standard) term as the key
synonym_groups = {
    "zero": ["zero", "none", "nil", "null"],
    "one": ["one", "won", "juan"],
    "two": ["two", "to", "too", "tu"],
    "three": ["three"],
    "four": ["four", "for", "fore"],
    "five": ["five"],
    "six": ["six"],
    "seven": ["seven"],
    "eight": ["eight", "ate"],
    "nine": ["nine"],
    "ten": ["ten"],
    "eleven": ["eleven"],
    "twelve": ["twelve", "dozen"],
    "never": ["never"],
    "half": ["half"],
    "once": ["once"]
}

Option 1: Normalize Synonyms to a Canonical Term

If you want to standardize your text (e.g., replace "won" with "one", "too" with "two"), create a reverse mapping from each synonym to its canonical term, then process the string:

# Build a case-insensitive map from synonym to canonical term
synonym_to_canonical = {}
for canonical, synonyms in synonym_groups.items():
    for syn in synonyms:
        synonym_to_canonical[syn.lower()] = canonical

def normalize_string(input_str):
    # Use regex to handle words with punctuation and preserve case
    import re
    def replace_word(match):
        word = match.group()
        lower_word = word.lower()
        if lower_word in synonym_to_canonical:
            # Match the original word's case
            if word.isupper():
                return synonym_to_canonical[lower_word].upper()
            elif word.istitle():
                return synonym_to_canonical[lower_word].title()
            else:
                return synonym_to_canonical[lower_word]
        return word
    return re.sub(r'\w+', replace_word, input_str)

# Example usage
input_text = "I won two tickets, but none of them are for the show at eight!"
print(normalize_string(input_text))
# Output: "I one two tickets, but zero of them are four the show at eight!"

Option 2: Expand Each Synonym to Its Full Group

If you want to replace each term with all related synonyms (e.g., replace "nil" with "zero, none, nil, null"), adjust the mapping to point to the full group string:

# Build a case-insensitive map from synonym to its full group
synonym_to_group = {}
for canonical, synonyms in synonym_groups.items():
    group_str = ", ".join(synonyms)
    for syn in synonyms:
        synonym_to_group[syn.lower()] = group_str

def expand_string(input_str):
    import re
    def replace_word(match):
        word = match.group()
        lower_word = word.lower()
        return synonym_to_group.get(lower_word, word)
    return re.sub(r'\w+', replace_word, input_str)

# Example usage
input_text = "I won two tickets, but none of them are for the show at eight!"
print(expand_string(input_text))
# Output: "I one, won, juan two, to, too, tu tickets, but zero, none, nil, null of them are four, for, fore the show at eight, ate!"

Key Notes:

  • Case Insensitivity: The code uses lower() to match synonyms regardless of how they're capitalized in the input, while preserving the original word's case in the output.
  • Punctuation Handling: Using regex (re.sub(r'\w+', ...)) ensures we match words even when they're attached to punctuation (like "won," or "none.").
  • Customization: You can tweak the synonym_groups dictionary to add more terms or adjust the canonical terms to fit your needs.

内容的提问来源于stack exchange,提问作者Jack

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:42:11