You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在buildDocWordMatrix内调用CleanText函数并传入其输出结果

代码修改方案

首先需要修复原两个函数的已知bug,再调整buildDocWordMatrix实现内部调用CleanText的需求:

1 修复CleanText函数的错误

原函数存在缩进错误,且返回值将所有文档的分词结果打平为一维列表,不符合词频矩阵的构建要求(需要保留每个文档的独立分词结果),修正后代码如下:

import nltk
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer

# 提前初始化依赖资源,remove_punc为你原有的去标点实现
stop = set(stopwords.words('english'))
def remove_punc(tokens):
    return [token for token in tokens if token.isalpha()]

def CleanText(doc_path) :
    # 修正缩进错误
    with open(doc_path, "r", encoding="utf-8") as myfile:
        corpus = myfile.read()
        docs = corpus.splitlines()
    doc_tokens = [nltk.word_tokenize(x) for x in docs] 
    doc_tokens_no_punc = [remove_punc(a_doc) for a_doc in doc_tokens]
    doc_tokens_no_punc_lower = [[element.lower() for element in x] for x in doc_tokens_no_punc]
    doc_tokens_clean = [[x for x in words if x not in stop] for words in doc_tokens_no_punc_lower]
    stemmer = PorterStemmer()
    # 修正打平逻辑,保留每个文档的分词层级
    doc_tokens_clean_stem = [[stemmer.stem(y) for y in x] for x in doc_tokens_clean]
    return doc_tokens_clean_stem

2 调整buildDocWordMatrix实现内部调用

原函数存在未定义变量x、向量追加位置缩进错误的问题,同时新增逻辑支持传入文件路径列表,内部自动调用CleanText完成预处理,修正后代码如下:

def buildDocWordMatrix(doc_path_list=None, preprocessed_doclist=None) :
    # 支持两种输入模式:直接传预处理好的分词列表,或传文件路径列表内部调用CleanText处理
    if preprocessed_doclist is None:
        if doc_path_list is None:
            raise ValueError("请传入文件路径列表doc_path_list或预处理后的分词列表preprocessed_doclist")
        # 内部调用CleanText处理所有输入文件
        preprocessed_doclist = []
        for path in doc_path_list:
            preprocessed_doclist.extend(CleanText(path))
    
    wordlist = []
    # 替换原未定义变量x为预处理后的文档列表
    for doc in preprocessed_doclist:
        for word in doc:
            if word not in wordlist:
                wordlist.append(word)
    docword= []
    for m in preprocessed_doclist:
        doc_vec = [0]*len(wordlist)
        for word in m:
            ind = wordlist.index(word)
            doc_vec[ind] += 1
        # 修正缩进:每个文档处理完再追加向量,避免重复添加
        docword.append(doc_vec)
    return docword, wordlist

使用示例

# 直接传入待处理的文件路径列表即可,函数内部自动完成预处理+矩阵构建
input_files = ["样本1.txt", "样本2.txt"]
doc_word_matrix, vocab = buildDocWordMatrix(doc_path_list=input_files)

内容的提问来源于stack exchange,提问作者user15227769

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 03:18:00