You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用NLTK查找特定concordance索引并关联展示结果

解决方案:关联索引与Concordance结果并实现单个索引查询

1. 基础准备:加载文本并初始化ConcordanceIndex

首先加载《白鲸记》文本,用NLTK的ConcordanceIndex类管理索引与上下文关联:

import nltk
from nltk.corpus import gutenberg
from nltk.text import ConcordanceIndex

# 首次运行需下载语料
nltk.download('gutenberg')

# 加载《白鲸记》单词列表
mobydick_tokens = gutenberg.words('melville-moby_dick.txt')
# 初始化ConcordanceIndex
ci = ConcordanceIndex(mobydick_tokens)

2. 输出带索引的Concordance结果

自定义函数实现默认concordance格式+索引绑定,解决大量结果难匹配的问题:

def concordance_with_indices(concordance_index, target_word, context_width=80, show_lines=25):
    # 获取目标词的所有出现索引
    offsets = concordance_index.offsets(target_word)
    if not offsets:
        print(f"未找到'{target_word}'的任何匹配")
        return

    print(f"「{target_word}」的共现上下文及索引(共{len(offsets)}条):")
    print("=" * (context_width + 20))
    
    # 遍历输出指定数量的结果
    for idx, offset in enumerate(offsets[:show_lines]):
        # 计算上下文左右边界
        left_start = max(0, offset - context_width // 2)
        right_end = min(len(concordance_index._tokens), offset + context_width // 2)
        
        # 拼接左右上下文
        left_context = ' '.join(concordance_index._tokens[left_start:offset])
        right_context = ' '.join(concordance_index._tokens[offset+1:right_end])
        
        # 格式化输出:左对齐索引,高亮目标词
        print(f"索引 {offset:6d}: ...{left_context.rjust(context_width//2)} *{target_word}* {right_context.ljust(context_width//2)}...")
    
    # 如果结果超过显示条数,提示剩余数量
    if len(offsets) > show_lines:
        print(f"\n... 仅展示前{show_lines}条,共{len(offsets)}条匹配")

# 调用示例:展示'monstrous'的前10条带索引的上下文
concordance_with_indices(ci, 'monstrous', context_width=80, show_lines=10)

3. 单个索引的查询与反向匹配

根据索引查询对应上下文

直接传入索引,快速定位该位置的词及上下文:

def get_context_by_index(tokens, target_index, context_width=80):
    if target_index < 0 or target_index >= len(tokens):
        print("索引超出文本范围")
        return
    
    left_start = max(0, target_index - context_width // 2)
    right_end = min(len(tokens), target_index + context_width // 2)
    
    left_context = ' '.join(tokens[left_start:target_index])
    target_word = tokens[target_index]
    right_context = ' '.join(tokens[target_index+1:right_end])
    
    print(f"索引 {target_index} 对应的上下文:")
    print(f"...{left_context.rjust(context_width//2)} *{target_word}* {right_context.ljust(context_width//2)}...")

# 调用示例:查询索引1301的上下文
get_context_by_index(mobydick_tokens, 1301)

根据上下文片段反向查找索引

如果记得部分上下文,可通过片段匹配找到对应的索引:

def find_indices_by_context_fragment(concordance_index, target_word, context_fragment):
    offsets = concordance_index.offsets(target_word)
    matching_indices = []
    
    for offset in offsets:
        # 取目标词前后各50个词的范围作为匹配上下文
        start = max(0, offset - 50)
        end = min(len(concordance_index._tokens), offset + 50)
        full_context = ' '.join(concordance_index._tokens[start:end])
        
        # 忽略大小写匹配上下文片段
        if context_fragment.lower() in full_context.lower():
            matching_indices.append(offset)
    
    if matching_indices:
        print(f"匹配上下文片段的索引:{matching_indices}")
    else:
        print("未找到匹配的索引")

# 调用示例:查找包含"size of it"的'monstrous'对应的索引
find_indices_by_context_fragment(ci, 'monstrous', 'size of it')

内容的提问来源于stack exchange,提问作者David Beales

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 18:27:42