You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TypeError问题求助:__init__()参数数量不匹配(Python文本分析)

解决TypeError: init() takes from 1 to 2 positional arguments but 4 were given错误

错误原因

错误出在CorpusLoader类的__init__方法中初始化KFold的代码。新版scikit-learn的KFold构造函数不再接受位置参数(样本数、折数、是否洗牌),而是使用关键字参数指定配置,直接传入三个位置参数会触发参数不匹配的错误。同时代码中还存在几处与新版sklearn API不兼容、变量名错误的问题。

修复步骤

  1. 修改KFold初始化方式:用关键字参数n_splits指定折数,shuffle控制是否洗牌,替代旧的位置参数写法。
  2. 修正shuffle参数赋值:__init__中self.shuffle应使用传入的参数,而非硬编码为True。
  3. 适配新版KFold索引生成:新版KFold需调用split()方法生成训练/测试索引,不能直接枚举实例。
  4. 更新KFold属性名:新版KFold用n_splits替代旧版的n_folds属性。
  5. 修正main函数变量名:loader = CorpusLoader(corpus, folds=12)中的corpus应为定义好的corpus4。
  6. 完善fileids调用逻辑:调用fileids()时必须指定train=True或test=True,否则会触发参数验证错误。

修改后的完整代码

from sklearn.model_selection import KFold


class CorpusLoader(object):
    """
    The corpus loader knows how to deal with an NLTK corpus at the top of a
    pipeline by simply taking as input a corpus to read from. It exposes both
    the data and the labels and can be set up to do cross-validation.
    If a number of folds is passed in for cross-validation, then the loader
    is smart about how to access data for train/test splits. Otherwise it will
    simply yield all documents in the corpus.
    """

    def __init__(self, corpus, folds=None, shuffle=True):
        self.n_docs = len(corpus.fileids())
        self.corpus = corpus
        self.folds  = folds
        self.shuffle = shuffle

        if folds is not None:
            self.folds = KFold(n_splits=folds, shuffle=shuffle)

    @property
    def n_folds(self):
        """
        Returns the number of folds if it exists; 0 otherwise.
        """
        if self.folds is None: return 0
        return self.folds.n_splits

    def fileids(self, fold=None, train=False, test=False):
        """
        Returns a listing of the documents filtering to retreive specific
        data from the folds/splits. If no fold, train, or test is specified
        then the method will return all fileids.
        If a fold is specified (should be an integer between 0 and folds),
        then the loader will return documents from that fold. Further, train
        or test must be specified to split the fold correctly.
        """
        if fold is None:
            return self.corpus.fileids()

        for fold_idx, (train_idx, test_idx) in enumerate(self.folds.split(range(self.n_docs))):
            if fold_idx == fold: break
        else:
            raise ValueError(
                "{} is not a fold, specify an integer less than {}".format(
                    fold, self.folds.n_splits
                )
            )

        if not (test or train) or (test and train):
            raise ValueError(
                "Please specify either train or test flag"
            )

        indices = train_idx if train else test_idx
        return [
            fileid for doc_idx, fileid in enumerate(self.corpus.fileids())
            if doc_idx in indices
        ]

    def labels(self, fold=None, train=False, test=False):
        """
        Fit will load a list of the labels from the corpus categories.
        If a fold is specified (should be an integer between 0 and folds),
        then the loader will return documents from that fold. Further, train
        or test must be specified to split the fold correctly.
        """
        return [
            self.corpus.categories(fileids=fileid)[0]
            for fileid in self.fileids(fold, train, test)
        ]

    def documents(self, fold=None, train=False, test=False):
        """
        A generator of documents being streamed from disk. Each document is
        a list of paragraphs, which are a list of sentences, which in turn is
        a list of tuples of (token, tag) pairs. All preprocessing is done by
        NLTK and the CorpusReader object this object wraps.
        If a fold is specified (should be an integer between 0 and folds),
        then the loader will return documents from that fold. Further, train
        or test must be specified to split the fold correctly. This method
        allows us to maintain the generator properties of document reads.
        """
        for fileid in self.fileids(fold, train, test):
            yield list(self.corpus.tagged(fileids=fileid))

if __name__ == '__main__':
    from reader import PickledCorpusReader

    corpus4 = PickledCorpusReader(nomi, r'.+\.txt')
    loader = CorpusLoader(corpus4, folds=12)

    for fid in loader.fileids(0, train=True):
        print(fid)

内容的提问来源于stack exchange,提问作者Andy Singal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 09:10:44