Python中如何合并多词汽车品牌识别结果的拆分项?
问题场景
我们有一个包含多词汽车品牌的制造商列表:
['Nissan', 'Honda', 'Rolls Royce', 'Ford', 'Kia', 'Land Rover']
用户输入句子:
I drive a Rolls Royce, a Ford, and a Land Rover.
需要识别句子中的所有汽车制造商,但当前用集合交集实现的代码会将多词品牌拆分为单个单词,返回结果为:
I found 5 automobiles in your sentence
Your automobiles are ['Royce', 'Rolls', 'Ford', 'Land', 'Rover']
期望输出是正确合并多词品牌的结果:
Your automobiles are ['Rolls Royce', 'Ford', 'Land Rover']
当前问题代码
import string result = [] autos = [ 'Mazda', 'Toyota', 'Volkswagon', 'Honda', 'Ford', 'Chevolet', 'Tesla', 'Kia', 'Hyndai', 'Rolls Royce', 'Isuzu', 'Jeep', 'Land Rover', 'Mitzubishi', 'Subaru' ] print("original autos",autos) ### Split each item new_autos = [] for auto in autos: new_autos.extend(auto.split()) print("\nnew_autos", new_autos) ### ask the user for a sentence sentence = input('Enter a sentence with favorite automobiles of your choice: ') title_case = sentence.title() ### remove punctuation new_sentence = title_case.translate(str.maketrans('', '', string.punctuation)) ### convert it to a list sentence_list = list(new_sentence.split(" ")) print(list(sentence_list)) set1 = set(sentence_list) set2 = set(new_autos) result = list(set1 & set2) ### tried differnt apporaches to getting the discovered items back in their original order ###result = list(set1.intersection(set2)) ###result = list(set1.intersection(set2)) ###result = sorted(set1.intersection(set2) ,key=lambda x:set1.index(x)) ####result = sorted(set1 & set2, key = set1.index) if len(result) > 0: substitution = sentence.split() substitution[sentence_list.index(result[-1])] = 'Brussels Sprouts' print(f"\n\nI found {len(result)} automobiles in your sentnece") print("Your automobiles are", result) print("Your final sentence with some brussels spouts is:", ' '.join(substitution)) else: print(f"\n\nI found no automobiles in your sentnece") print("The automobiles found is an empty list:", result) print("Your final sentence is:", sentence)
问题分析
当前代码的核心问题是将所有汽车品牌拆分为单个单词存入集合,丢失了多词品牌的上下文关联。当匹配时,只能识别出单个单词,无法还原成原始的多词品牌名称。
修复方案
核心思路是优先匹配多词品牌,避免拆分后丢失上下文。具体步骤:
- 将汽车品牌列表按单词数量降序排序,确保多词品牌被优先匹配
- 处理输入句子,去除标点并保留可匹配的格式
- 遍历排序后的品牌,检查是否存在于处理后的句子中,存在则加入结果,并从句子中移除该品牌(防止重复匹配)
修复后的代码
import string autos = [ 'Mazda', 'Toyota', 'Volkswagon', 'Honda', 'Ford', 'Chevolet', 'Tesla', 'Kia', 'Hyndai', 'Rolls Royce', 'Isuzu', 'Jeep', 'Land Rover', 'Mitzubishi', 'Subaru' ] print("Original autos:", autos) # 按品牌的单词数量降序排序,优先匹配多词品牌 autos_sorted = sorted(autos, key=lambda x: len(x.split()), reverse=True) # 获取用户输入并处理 sentence = input('Enter a sentence with your favorite automobiles: ') # 去除标点,保留原始大小写(或统一为标题格式,确保匹配) processed_sentence = sentence.translate(str.maketrans('', '', string.punctuation)) found_autos = [] temp_sentence = processed_sentence # 临时句子,用于移除已匹配的品牌 for auto in autos_sorted: if auto in temp_sentence: found_autos.append(auto) # 从临时句子中移除已匹配的品牌,避免重复匹配 temp_sentence = temp_sentence.replace(auto, '') # 保持结果在原句子中的出现顺序 def get_order_in_sentence(item): return processed_sentence.index(item) found_autos_sorted = sorted(found_autos, key=get_order_in_sentence) # 输出结果 if found_autos_sorted: print(f"\nI found {len(found_autos_sorted)} automobiles in your sentence") print("Your automobiles are:", found_autos_sorted) # 替换最后一个品牌为 Brussels Sprouts(保留原功能) original_words = sentence.split() # 找到最后一个品牌在原句子中的位置 last_auto = found_autos_sorted[-1] # 处理原句子中的标点,找到对应单词位置 for i, word in enumerate(original_words): # 去除单词标点后匹配 clean_word = word.translate(str.maketrans('', '', string.punctuation)) if last_auto.startswith(clean_word) and last_auto in processed_sentence: original_words[i] = 'Brussels Sprouts' # 如果是多词品牌,后续单词也替换为空(可选,这里简化处理) if len(last_auto.split()) > 1: for j in range(i+1, len(original_words)): clean_j = original_words[j].translate(str.maketrans('', '', string.punctuation)) if clean_j in last_auto.split(): original_words[j] = '' original_words = [w for w in original_words if w] break print("Your final sentence with some brussels sprouts is:", ' '.join(original_words)) else: print("\nI found no automobiles in your sentence") print("The automobiles found is an empty list:", found_autos_sorted) print("Your final sentence is:", sentence)
最终效果
输入句子I drive a Rolls Royce, a Ford, and a Land Rover.后,输出:
Original autos: ['Mazda', 'Toyota', 'Volkswagon', 'Honda', 'Ford', 'Chevolet', 'Tesla', 'Kia', 'Hyndai', 'Rolls Royce', 'Isuzu', 'Jeep', 'Land Rover', 'Mitzubishi', 'Subaru']
Enter a sentence with your favorite automobiles: I drive a Rolls Royce, a Ford, and a Land Rover.I found 3 automobiles in your sentence
Your automobiles are: ['Rolls Royce', 'Ford', 'Land Rover']
Your final sentence with some brussels sprouts is: I drive a Rolls Royce, a Ford, and a Brussels Sprouts.
内容的提问来源于stack exchange,提问作者Scott Bing

