如何以Pythonic方式生成含斜杠的英文句子的所有变体?
问题描述
给定包含正斜杠(表示“或”)的英文句子,需生成所有可能的变体字符串:每个斜杠分隔的选项需替换为独立句子,句子可包含任意数量斜杠或无斜杠。
示例
示例1
输入:
example = "bear a/little/no resemblance to sth/sb/whatever"
输出:
alternatives = ['bear a resemblance to sth', 'bear a resemblance to sb', 'bear a resemblance to whatever', 'bear little resemblance to sth', 'bear little resemblance to sb', 'bear little resemblance to whatever', 'bear no resemblance to sth', 'bear no resemblance to sb', 'bear no resemblance to whatever']
示例2
输入:
example = "beat about/around the bush"
输出:
alternatives = ['beat about the bush', 'beat around the bush']
示例3
输入:
example = "become available/rich/a writer, etc."
输出:
alternatives = ['become available', 'become rich', 'become a writer']
现有非Pythonic实现
用户提供的原始代码通过多次正则匹配、手动拆分替换、循环清理列表的方式实现,逻辑冗余且不够简洁:
alt = [] #short for alternatives # multiple cleaning stages for every baseword ##1## remove ', etc.' a = example.replace(', etc.', '') ##2## does this string have / / in it regex = re.compile(r'(\w+)\/(\w+)\/(\w+)') match = regex.search(a) delete_later = [] #a list with sentence to delete later from alt as it cleans up old used sentences if match: part = a.partition(match.group(0)) s1 = part[0]+match.group(1)+part[2] s2 = part[0]+match.group(2)+part[2] s3 = part[0]+match.group(3)+part[2] alt.append(s1) alt.append(s2) alt.append(s3) #check again: for _ in range(10): for item in alt: regex = re.compile(r'(\w+)\/(\w+)\/(\w+)') match = regex.search(item) if match: delete_later.append(item) part = item.partition(match.group(0)) s1 = part[0]+match.group(1)+part[2] s2 = part[0]+match.group(2)+part[2] s3 = part[0]+match.group(3)+part[2] alt.append(s1) alt.append(s2) alt.append(s3) #clean up for i in delete_later: try: #avoid Traceback: ValueError: list.remove(x): x not in list alt.remove(i) except: pass ##3## does this string have / in it if len(alt) > 0: for _ in range(10): for item in alt: regex = re.compile(r'(\w+)\/(\w+)') match = regex.search(item) if match: delete_later.append(item) part = item.partition(match.group(0)) s1 = part[0]+match.group(1)+part[2] s2 = part[0]+match.group(2)+part[2] alt.append(s1) alt.append(s2) #clean up for i in delete_later: try: #avoid Traceback: ValueError: list.remove(x): x not in list alt.remove(i) except: pass #else: #check for the 1st time regex = re.compile(r'(\w+)\/(\w+)') match = regex.search(a) delete_later = [] #a list with sentence to delete later from alt as it cleans up old used sentences if match: part = a.partition(match.group(0)) s1 = part[0]+match.group(1)+part[2] s2 = part[0]+match.group(2)+part[2] alt.append(s1) alt.append(s2) #check again: for _ in range(10): for item in alt: regex = re.compile(r'(\w+)\/(\w+)') match = regex.search(item) if match: delete_later.append(item) part = item.partition(match.group(0)) s1 = part[0]+match.group(1)+part[2] s2 = part[0]+match.group(2)+part[2] alt.append(s1) alt.append(s2) #clean up for i in delete_later: try: #avoid Traceback: ValueError: list.remove(x): x not in list alt.remove(i) except: pass for i,e in enumerate(alt , 1): print(i,e)
Pythonic优雅实现
我们可以利用正则拆分出所有固定片段和可选分支,再通过itertools.product计算笛卡尔积来生成所有组合,逻辑简洁且高效:
import re from itertools import product def generate_variants(sentence): # 清理输入:去掉', etc.'后缀 cleaned = sentence.replace(', etc.', '') # 正则拆分:分离固定文本段和带斜杠的可选组 parts = re.split(r'([^\s/]+(?:/[^\s/]+)+)', cleaned) # 构建选项列表:可选组分拆为选项,固定文本保持单元素列表 options = [] for part in parts: if '/' in part: options.append(part.split('/')) else: options.append([part]) # 计算笛卡尔积并拼接所有组合 return [''.join(combo).strip() for combo in product(*options)] # 测试示例 if __name__ == "__main__": example1 = "bear a/little/no resemblance to sth/sb/whatever" print(generate_variants(example1)) example2 = "beat about/around the bush" print(generate_variants(example2)) example3 = "become available/rich/a writer, etc." print(generate_variants(example3))
代码说明
- 输入清理:先移除
, etc.这类无关后缀; - 拆分句子:用正则把句子拆分成固定文本段和带斜杠的可选组;
- 构建选项列表:将每个可选组拆分成选项列表,固定文本段保持为单元素列表;
- 生成所有组合:用
itertools.product计算所有选项的笛卡尔积,拼接后得到所有变体字符串。
这个实现无需循环清理列表,逻辑清晰,能处理任意数量的斜杠分支,完全符合需求且更符合Python风格。
内容的提问来源于stack exchange,提问作者ZZZ
相关产品推荐
相关产品推荐

