如何对字符串中的特定词汇进行扩展?附待扩展词汇列表
Great question! Let's walk through how to work with your synonym lists to either normalize (standardize) or expand terms in a string. I'll use Python for examples since it's widely used for text processing, but the logic translates to other languages too.
First, let's restructure your separate lists into a more manageable dictionary—this will make it easier to map synonyms to their target terms or groups:
# Organize your synonym groups with a canonical (standard) term as the key synonym_groups = { "zero": ["zero", "none", "nil", "null"], "one": ["one", "won", "juan"], "two": ["two", "to", "too", "tu"], "three": ["three"], "four": ["four", "for", "fore"], "five": ["five"], "six": ["six"], "seven": ["seven"], "eight": ["eight", "ate"], "nine": ["nine"], "ten": ["ten"], "eleven": ["eleven"], "twelve": ["twelve", "dozen"], "never": ["never"], "half": ["half"], "once": ["once"] }
Option 1: Normalize Synonyms to a Canonical Term
If you want to standardize your text (e.g., replace "won" with "one", "too" with "two"), create a reverse mapping from each synonym to its canonical term, then process the string:
# Build a case-insensitive map from synonym to canonical term synonym_to_canonical = {} for canonical, synonyms in synonym_groups.items(): for syn in synonyms: synonym_to_canonical[syn.lower()] = canonical def normalize_string(input_str): # Use regex to handle words with punctuation and preserve case import re def replace_word(match): word = match.group() lower_word = word.lower() if lower_word in synonym_to_canonical: # Match the original word's case if word.isupper(): return synonym_to_canonical[lower_word].upper() elif word.istitle(): return synonym_to_canonical[lower_word].title() else: return synonym_to_canonical[lower_word] return word return re.sub(r'\w+', replace_word, input_str) # Example usage input_text = "I won two tickets, but none of them are for the show at eight!" print(normalize_string(input_text)) # Output: "I one two tickets, but zero of them are four the show at eight!"
Option 2: Expand Each Synonym to Its Full Group
If you want to replace each term with all related synonyms (e.g., replace "nil" with "zero, none, nil, null"), adjust the mapping to point to the full group string:
# Build a case-insensitive map from synonym to its full group synonym_to_group = {} for canonical, synonyms in synonym_groups.items(): group_str = ", ".join(synonyms) for syn in synonyms: synonym_to_group[syn.lower()] = group_str def expand_string(input_str): import re def replace_word(match): word = match.group() lower_word = word.lower() return synonym_to_group.get(lower_word, word) return re.sub(r'\w+', replace_word, input_str) # Example usage input_text = "I won two tickets, but none of them are for the show at eight!" print(expand_string(input_text)) # Output: "I one, won, juan two, to, too, tu tickets, but zero, none, nil, null of them are four, for, fore the show at eight, ate!"
Key Notes:
- Case Insensitivity: The code uses
lower()to match synonyms regardless of how they're capitalized in the input, while preserving the original word's case in the output. - Punctuation Handling: Using regex (
re.sub(r'\w+', ...)) ensures we match words even when they're attached to punctuation (like "won," or "none."). - Customization: You can tweak the
synonym_groupsdictionary to add more terms or adjust the canonical terms to fit your needs.
内容的提问来源于stack exchange,提问作者Jack

