You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python构建信息检索系统时索引环节路径参数错误求助

问题分析与解决

错误原因

报错的直接原因是folder_path = documents这一行中,documents是列表类型,但os.listdir()函数要求传入的参数必须是字符串、字节、路径对象等,不能是列表。

修正方案

把folder_path赋值为你的文档所在文件夹的字符串路径(绝对路径或相对路径都可以),比如你的文档都放在当前目录下的docs_folder文件夹里,就这么写:

folder_path = "./docs_folder"  # 替换成你的实际文件夹路径

修正后的完整代码:

import os
import nltk
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

# Step 2: Collect the dataset
folder_path = "./docs_folder"  # 这里替换成你的文档文件夹的实际路径
docs = []
# 过滤非文本文件,避免读取无效文件
for file_name in os.listdir(folder_path):
    file_path = os.path.join(folder_path, file_name)
    # 只处理文件,跳过文件夹,可选过滤.txt后缀
    if os.path.isfile(file_path) and file_name.endswith(".txt"):
        try:
            with open(file_path, "r", encoding="utf-8") as f:
                docs.append(f.read())
        except UnicodeDecodeError:
            print(f"无法读取文件: {file_name},编码异常")

# Step 3: Indexing
nltk.download('stopwords')
nltk.download('wordnet')
stop_words = set(stopwords.words('english'))
lemmatizer = WordNetLemmatizer()

def tokenize_and_lemmatize(text):
    tokens = nltk.word_tokenize(text.lower())
    tokens = [lemmatizer.lemmatize(token) for token in tokens if token.isalpha()]
    tokens = [token for token in tokens if token not in stop_words]
    return tokens

# Step 4: Tf-idf retrieval model
# 新版本sklearn中,自定义tokenizer需配合token_pattern=None,避免默认规则干扰
vectorizer = TfidfVectorizer(tokenizer=tokenize_and_lemmatize, token_pattern=None)
tfidf_matrix = vectorizer.fit_transform(docs)

# Step 5: Query
query = "machine learning"
query_vec = vectorizer.transform([query])

# Step 6: Similarity measure
similarity_scores = cosine_similarity(tfidf_matrix, query_vec)

# Step 7: Display results
results = [(score[0], doc) for score, doc in zip(similarity_scores, docs)]
results = sorted(results, reverse=True)

for score, doc in results:
    print(f"{score:.3f}: {doc[:50]}...")

额外优化说明

  • 增加文件过滤逻辑,只读取.txt文件,跳过文件夹和非文本文件,避免无效读取
  • 加入编码处理与异常捕获,防止因文件编码问题导致程序崩溃
  • 给TfidfVectorizer添加token_pattern=None,避免默认分词规则干扰自定义分词结果

内容的提问来源于stack exchange,提问作者Bayan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 08:24:58