You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python基于DataFrame构建倒排索引(Inverted Index)

倒排索引实现方案

以下是基于Python+pandas的实现代码,可直接适配你的DataFrame结构:

前置依赖

需要用到的基础库,无需额外安装第三方包:

  • pandas:用来读取/处理你的DataFrame
  • collections.defaultdict:用来初始化倒排索引结构

完整实现代码

import pandas as pd
from collections import defaultdict

# --------------- 1. 加载你的DataFrame,此处为示例构造,替换为你自己的数据源即可 ---------------
df = pd.DataFrame({
    'document': ['Ancient Egypt', 'Nile River'],
    'content': [
        'Ancient Egypt was a civilization of ancient North Africa',
        'The Nile is a major north flowing river in northeastern Africa'
    ]
})

# --------------- 2. 核心倒排索引构建逻辑 ---------------
inverted_index = defaultdict(list)

for _, row in df.iterrows():
    doc = row['document']
    # 按空格拆分内容为单词,可根据需求添加预处理逻辑
    words = row['content'].split()
    # 对单篇文档内的单词去重,避免同一文档重复录入
    for word in set(words):
        if doc not in inverted_index[word]:
            inverted_index[word].append(doc)

# 转为普通字典,完全匹配你需要的输出格式
result = dict(inverted_index)

输出验证

打印result即可得到你要的格式:

print(result)
# 输出示例:
# {'a': ['Ancient Egypt', 'Nile River'],
#  'Egypt': ['Ancient Egypt'],
#  'is': ['Nile River']}

可选优化项

如果需要优化索引准确性,可以添加单词预处理逻辑:

  • 统一单词大小写:word = word.lower(),避免A和a被识别为两个不同单词
  • 过滤标点符号:导入string库后用word = word.strip(string.punctuation),去除单词两端的逗号、句号等符号
  • 过滤停用词:可自定义停用词列表,过滤掉the、of这类无实际检索意义的单词

内容的提问来源于stack exchange,提问作者N-Sam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 03:54:05