如何基于各位置词列表生成字符串的所有组合
问题描述
我有如下字符串:
original_text = "womens wear apparel bike"
现在,original_text中的每个词对应一组备选词,如下所示:
text_to_generate = [['females', 'ladies'], 'wear', ['clothing', 'clothes'], ['biking', 'cycling', 'running']]
我需要利用该列表中的词生成所有可能的短语,期望输出示例如下:
text1 = 'females wear clothing biking' text2 = 'females wear clothes cycling' text3 = 'ladies wear clothing biking' text4 = 'ladies wear clothes cycling' text5 = 'ladies wear clothes running'
各位置的词列表长度可能不一致。以下是我目前的尝试代码:
original_text = "womens wear apparel bike" alternates_dict = { "mens": ["males"], "vitamins": ["supplements"], "womens": ["females", "ladies"], "shoes": ["footwear"], "apparel": ["clothing", "clothes"], "kids": ["childrens", "childs"], "motorcycle": ["motorbike"], "watercraft": ["boat"], "medicine": ["medication"], "supplements": ["vitamins"], "t-shirt": ["shirt"], "pram": ["stroller"], "bike": ["biking", "cycling"], } splitted = original_text.split() for i in range(0,len(splitted)): if splitted[i] in alternates_dict.keys(): splitted[i] = alternates_dict[splitted[i]] for word in splitted[i]: update = original_text.replace(original_text.split()[i], word) print(update) print(splitted)
解决方案
你的当前代码只能逐个替换单个位置的备选词,无法生成所有位置的组合短语。要实现需求,可借助Python标准库itertools中的product函数——它能直接生成多个可迭代对象的笛卡尔积,完美匹配所有短语组合的生成逻辑。
实现步骤
- 将原始文本的每个词转换为对应备选词列表:若词在字典中有备选则用备选列表,无备选则将原词包装为单元素列表(确保每个位置都是可迭代对象)。
- 用
itertools.product生成所有位置词的笛卡尔积。 - 将每个组合中的词用空格拼接成完整短语。
完整代码
import itertools original_text = "womens wear apparel bike" alternates_dict = { "mens": ["males"], "vitamins": ["supplements"], "womens": ["females", "ladies"], "shoes": ["footwear"], "apparel": ["clothing", "clothes"], "kids": ["childrens", "childs"], "motorcycle": ["motorbike"], "watercraft": ["boat"], "medicine": ["medication"], "supplements": ["vitamins"], "t-shirt": ["shirt"], "pram": ["stroller"], "bike": ["biking", "cycling", "running"], # 补充running以匹配示例需求 } # 处理每个词,生成对应的备选选项列表 word_options = [] for word in original_text.split(): word_options.append(alternates_dict.get(word, [word])) # 生成所有组合并按格式输出 for idx, combo in enumerate(itertools.product(*word_options), start=1): phrase = ' '.join(combo) print(f'text{idx} = \'{phrase}\'')
输出结果
text1 = 'females wear clothing biking' text2 = 'females wear clothing cycling' text3 = 'females wear clothing running' text4 = 'females wear clothes biking' text5 = 'females wear clothes cycling' text6 = 'females wear clothes running' text7 = 'ladies wear clothing biking' text8 = 'ladies wear clothing cycling' text9 = 'ladies wear clothing running' text10 = 'ladies wear clothes biking' text11 = 'ladies wear clothes cycling' text12 = 'ladies wear clothes running'
内容的提问来源于stack exchange,提问作者SamuraiSam
相关产品推荐
相关产品推荐

