如何用NLTK查找特定concordance索引并关联展示结果
解决方案:关联索引与Concordance结果并实现单个索引查询
1. 基础准备:加载文本并初始化ConcordanceIndex
首先加载《白鲸记》文本,用NLTK的ConcordanceIndex类管理索引与上下文关联:
import nltk from nltk.corpus import gutenberg from nltk.text import ConcordanceIndex # 首次运行需下载语料 nltk.download('gutenberg') # 加载《白鲸记》单词列表 mobydick_tokens = gutenberg.words('melville-moby_dick.txt') # 初始化ConcordanceIndex ci = ConcordanceIndex(mobydick_tokens)
2. 输出带索引的Concordance结果
自定义函数实现默认concordance格式+索引绑定,解决大量结果难匹配的问题:
def concordance_with_indices(concordance_index, target_word, context_width=80, show_lines=25): # 获取目标词的所有出现索引 offsets = concordance_index.offsets(target_word) if not offsets: print(f"未找到'{target_word}'的任何匹配") return print(f"「{target_word}」的共现上下文及索引(共{len(offsets)}条):") print("=" * (context_width + 20)) # 遍历输出指定数量的结果 for idx, offset in enumerate(offsets[:show_lines]): # 计算上下文左右边界 left_start = max(0, offset - context_width // 2) right_end = min(len(concordance_index._tokens), offset + context_width // 2) # 拼接左右上下文 left_context = ' '.join(concordance_index._tokens[left_start:offset]) right_context = ' '.join(concordance_index._tokens[offset+1:right_end]) # 格式化输出:左对齐索引,高亮目标词 print(f"索引 {offset:6d}: ...{left_context.rjust(context_width//2)} *{target_word}* {right_context.ljust(context_width//2)}...") # 如果结果超过显示条数,提示剩余数量 if len(offsets) > show_lines: print(f"\n... 仅展示前{show_lines}条,共{len(offsets)}条匹配") # 调用示例:展示'monstrous'的前10条带索引的上下文 concordance_with_indices(ci, 'monstrous', context_width=80, show_lines=10)
3. 单个索引的查询与反向匹配
根据索引查询对应上下文
直接传入索引,快速定位该位置的词及上下文:
def get_context_by_index(tokens, target_index, context_width=80): if target_index < 0 or target_index >= len(tokens): print("索引超出文本范围") return left_start = max(0, target_index - context_width // 2) right_end = min(len(tokens), target_index + context_width // 2) left_context = ' '.join(tokens[left_start:target_index]) target_word = tokens[target_index] right_context = ' '.join(tokens[target_index+1:right_end]) print(f"索引 {target_index} 对应的上下文:") print(f"...{left_context.rjust(context_width//2)} *{target_word}* {right_context.ljust(context_width//2)}...") # 调用示例:查询索引1301的上下文 get_context_by_index(mobydick_tokens, 1301)
根据上下文片段反向查找索引
如果记得部分上下文,可通过片段匹配找到对应的索引:
def find_indices_by_context_fragment(concordance_index, target_word, context_fragment): offsets = concordance_index.offsets(target_word) matching_indices = [] for offset in offsets: # 取目标词前后各50个词的范围作为匹配上下文 start = max(0, offset - 50) end = min(len(concordance_index._tokens), offset + 50) full_context = ' '.join(concordance_index._tokens[start:end]) # 忽略大小写匹配上下文片段 if context_fragment.lower() in full_context.lower(): matching_indices.append(offset) if matching_indices: print(f"匹配上下文片段的索引:{matching_indices}") else: print("未找到匹配的索引") # 调用示例:查找包含"size of it"的'monstrous'对应的索引 find_indices_by_context_fragment(ci, 'monstrous', 'size of it')
内容的提问来源于stack exchange,提问作者David Beales
相关产品推荐
相关产品推荐

