Python如何遍历列表所有索引位置提取双语对齐词汇
问题原因
你出现报错的核心原因是:hyphen_split 本身是嵌套列表,存储了当前句对所有的对齐索引对(结构类似 [["0","0"], ["1","1"], ...]),你用 hyphen_split[0:] 是把所有索引对整体提取为一个新列表,此时 align_index[0] 是第一个索引对的子列表,而非可转为int的索引字符串,因此触发类型错误。
之前的代码只取了hyphen_split[0]也就是第一个索引对,所以只能输出第一个对齐词汇,要输出所有对齐项,只需要加一层循环遍历所有索引对即可。
修复后代码
en_de = [ [['The', 'hat', 'is', 'on', 'the', 'table', '.'], ['Der', 'Hut', 'liegt', 'auf', 'dem', 'Tisch', '.'], '0-0 1-1 2-2 3-3 4-4 5-5 6-6'], [['The', 'picture', 'is', 'on', 'the', 'wall', '.'], ['Das', 'Bild', 'hängt', 'an', 'der', 'Wand', '.'], '0-0 1-1 2-2 3-3 4-4 5-5 6-6'], [['The', 'bottle', 'is', 'under', 'the', 'sink', '.'], ['Die', 'Flasche', 'ist', 'under', 'dem', 'Waschbecken', '.'], '0-0 1-1 2-2 3-3 4-4 5-5 6-6'] ] for group in en_de: src_sent = group[0] tgt_sent = group[1] aligns = group[2] # 拆分所有对齐对 split_aligns = aligns.split() hyphen_split = [align.split("-") for align in split_aligns] # 遍历每一个对齐对输出 for src_idx, tgt_idx in hyphen_split: print(f"{src_sent[int(src_idx)]} - {tgt_sent[int(tgt_idx)]}") # 不同句对之间加空行分隔,可根据需要删除 print()
运行输出示例
The - Der hat - Hut is - liegt on - auf the - dem table - Tisch . - . The - Das picture - Bild is - hängt on - an the - der wall - Wand . - . The - Die bottle - Flasche is - ist under - under the - dem sink - Waschbecken . - .
内容的提问来源于stack exchange,提问作者c_metaphorique
相关产品推荐
相关产品推荐

