TypeError问题求助:__init__()参数数量不匹配(Python文本分析)
解决TypeError: init() takes from 1 to 2 positional arguments but 4 were given错误
错误原因
错误出在CorpusLoader类的__init__方法中初始化KFold的代码。新版scikit-learn的KFold构造函数不再接受位置参数(样本数、折数、是否洗牌),而是使用关键字参数指定配置,直接传入三个位置参数会触发参数不匹配的错误。同时代码中还存在几处与新版sklearn API不兼容、变量名错误的问题。
修复步骤
- 修改KFold初始化方式:用关键字参数
n_splits指定折数,shuffle控制是否洗牌,替代旧的位置参数写法。 - 修正shuffle参数赋值:
__init__中self.shuffle应使用传入的参数,而非硬编码为True。 - 适配新版KFold索引生成:新版
KFold需调用split()方法生成训练/测试索引,不能直接枚举实例。 - 更新KFold属性名:新版
KFold用n_splits替代旧版的n_folds属性。 - 修正main函数变量名:
loader = CorpusLoader(corpus, folds=12)中的corpus应为定义好的corpus4。 - 完善fileids调用逻辑:调用
fileids()时必须指定train=True或test=True,否则会触发参数验证错误。
修改后的完整代码
from sklearn.model_selection import KFold class CorpusLoader(object): """ The corpus loader knows how to deal with an NLTK corpus at the top of a pipeline by simply taking as input a corpus to read from. It exposes both the data and the labels and can be set up to do cross-validation. If a number of folds is passed in for cross-validation, then the loader is smart about how to access data for train/test splits. Otherwise it will simply yield all documents in the corpus. """ def __init__(self, corpus, folds=None, shuffle=True): self.n_docs = len(corpus.fileids()) self.corpus = corpus self.folds = folds self.shuffle = shuffle if folds is not None: self.folds = KFold(n_splits=folds, shuffle=shuffle) @property def n_folds(self): """ Returns the number of folds if it exists; 0 otherwise. """ if self.folds is None: return 0 return self.folds.n_splits def fileids(self, fold=None, train=False, test=False): """ Returns a listing of the documents filtering to retreive specific data from the folds/splits. If no fold, train, or test is specified then the method will return all fileids. If a fold is specified (should be an integer between 0 and folds), then the loader will return documents from that fold. Further, train or test must be specified to split the fold correctly. """ if fold is None: return self.corpus.fileids() for fold_idx, (train_idx, test_idx) in enumerate(self.folds.split(range(self.n_docs))): if fold_idx == fold: break else: raise ValueError( "{} is not a fold, specify an integer less than {}".format( fold, self.folds.n_splits ) ) if not (test or train) or (test and train): raise ValueError( "Please specify either train or test flag" ) indices = train_idx if train else test_idx return [ fileid for doc_idx, fileid in enumerate(self.corpus.fileids()) if doc_idx in indices ] def labels(self, fold=None, train=False, test=False): """ Fit will load a list of the labels from the corpus categories. If a fold is specified (should be an integer between 0 and folds), then the loader will return documents from that fold. Further, train or test must be specified to split the fold correctly. """ return [ self.corpus.categories(fileids=fileid)[0] for fileid in self.fileids(fold, train, test) ] def documents(self, fold=None, train=False, test=False): """ A generator of documents being streamed from disk. Each document is a list of paragraphs, which are a list of sentences, which in turn is a list of tuples of (token, tag) pairs. All preprocessing is done by NLTK and the CorpusReader object this object wraps. If a fold is specified (should be an integer between 0 and folds), then the loader will return documents from that fold. Further, train or test must be specified to split the fold correctly. This method allows us to maintain the generator properties of document reads. """ for fileid in self.fileids(fold, train, test): yield list(self.corpus.tagged(fileids=fileid)) if __name__ == '__main__': from reader import PickledCorpusReader corpus4 = PickledCorpusReader(nomi, r'.+\.txt') loader = CorpusLoader(corpus4, folds=12) for fid in loader.fileids(0, train=True): print(fid)
内容的提问来源于stack exchange,提问作者Andy Singal
相关产品推荐
相关产品推荐

