如何使用Python基于DataFrame构建倒排索引(Inverted Index)
倒排索引实现方案
以下是基于Python+pandas的实现代码,可直接适配你的DataFrame结构:
前置依赖
需要用到的基础库,无需额外安装第三方包:
- pandas:用来读取/处理你的DataFrame
- collections.defaultdict:用来初始化倒排索引结构
完整实现代码
import pandas as pd from collections import defaultdict # --------------- 1. 加载你的DataFrame,此处为示例构造,替换为你自己的数据源即可 --------------- df = pd.DataFrame({ 'document': ['Ancient Egypt', 'Nile River'], 'content': [ 'Ancient Egypt was a civilization of ancient North Africa', 'The Nile is a major north flowing river in northeastern Africa' ] }) # --------------- 2. 核心倒排索引构建逻辑 --------------- inverted_index = defaultdict(list) for _, row in df.iterrows(): doc = row['document'] # 按空格拆分内容为单词,可根据需求添加预处理逻辑 words = row['content'].split() # 对单篇文档内的单词去重,避免同一文档重复录入 for word in set(words): if doc not in inverted_index[word]: inverted_index[word].append(doc) # 转为普通字典,完全匹配你需要的输出格式 result = dict(inverted_index)
输出验证
打印result即可得到你要的格式:
print(result) # 输出示例: # {'a': ['Ancient Egypt', 'Nile River'], # 'Egypt': ['Ancient Egypt'], # 'is': ['Nile River']}
可选优化项
如果需要优化索引准确性,可以添加单词预处理逻辑:
- 统一单词大小写:
word = word.lower(),避免A和a被识别为两个不同单词 - 过滤标点符号:导入
string库后用word = word.strip(string.punctuation),去除单词两端的逗号、句号等符号 - 过滤停用词:可自定义停用词列表,过滤掉
the、of这类无实际检索意义的单词
内容的提问来源于stack exchange,提问作者N-Sam
相关产品推荐
相关产品推荐

