You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从TXT文件创建defaultdict?解决大文件索引越界报错

基于两个TXT文件构建TF-IDF嵌套defaultdict的问题

需求说明

需要利用两个TXT文件的数据创建结构为 {textnum : {word : tf*idf}} 的defaultdict:

  • 第一个TXT文件格式为word idf,示例数据:
acceler             4.634728988229636 
accept              2.32949254591397 
access              3.0633909220278057 
accid               3.9512437185814275 
acclaim             4.634728988229636 
  • 第二个TXT文件格式为textnum word tf,示例数据:
0097       about        0.07894736842105263 
0097        abus        0.02631578947368421 
0098      acceler       0.02631578947368421 
0098       across       0.02631578947368421 
0099      admonish      0.02631578947368421 
0099       after        0.05263157894736842

可用库推荐

  • collections:自带的defaultdict可直接实现嵌套字典结构,无需额外依赖
  • pandas:处理大文件更高效,支持快速读取、合并数据并批量计算TF-IDF,适合大规模数据场景
  • numpy:配合完成数值型的TF-IDF乘积运算,提升计算效率

代码问题与报错

编写的代码如下:

from collections import defaultdict
tf_idf_dict = defaultdict(dict) 

def read_in(path):
    with open(path, "r") as r:
        l = r.readlines()
        return l


def tf_idf_calc(tf_file, idf_file, d):
    tf_line = [[item for item in line.split()] for line in read_in(tf_file)]
    idf_line = [[item for item in line.split()] for line in read_in(idf_file)]
    for line in tf_line:
        for row in idf_line:
            if line[1] == row[0]:
                d[line[0]][line[1]] = float(line[2]) * float(row[1])

    return d


def nested_to_defaultdict(d): #converts the nested dictionary to defaultdict
    if not isinstance(d, dict):
        return d
    return defaultdict(lambda: 0, {key: nested_to_defaultdict(value) for key, value in d.items()})

代码处理小文件正常,但处理大文件时触发报错:

line 15, in tf_idf_calc
if line[1] == row[0]:
IndexError: list index out of range

报错翻译:

第15行,在tf_idf_calc函数中
if line[1] == row[0]:
索引错误:列表索引超出范围

问题分析与修复方案

  1. 报错原因:大文件中存在空行或格式异常的行,调用split()后得到空列表或元素数量不足的列表,导致访问line[1]或row[0]时触发索引越界;同时原代码双重循环遍历TF和IDF数据,时间复杂度为O(M*N),大文件下效率极低。

  2. 修复方案:

    • 过滤无效行:读取文件时跳过空行,检查每行拆分后的元素数量是否符合格式要求
    • 优化匹配逻辑:先将IDF数据转为字典,避免双重循环,将时间复杂度降至O(M+N)

修复后的代码示例:

from collections import defaultdict

def load_idf_dict(idf_file):
    idf_dict = {}
    with open(idf_file, "r") as f:
        for line in f:
            line = line.strip()
            if not line:
                continue
            parts = line.split()
            if len(parts) != 2:
                continue
            word, idf = parts
            idf_dict[word] = float(idf)
    return idf_dict

def build_tfidf_defaultdict(tf_file, idf_file):
    idf_dict = load_idf_dict(idf_file)
    tfidf_dict = defaultdict(dict)
    
    with open(tf_file, "r") as f:
        for line in f:
            line = line.strip()
            if not line:
                continue
            parts = line.split()
            if len(parts) != 3:
                continue
            textnum, word, tf = parts
            if word in idf_dict:
                tfidf_dict[textnum][word] = float(tf) * idf_dict[word]
    
    # 转换为嵌套defaultdict(按需使用)
    nested_default = defaultdict(lambda: defaultdict(lambda: 0))
    for textnum, word_dict in tfidf_dict.items():
        nested_default[textnum].update(word_dict)
    return nested_default

# 使用示例
result = build_tfidf_defaultdict("tf_file.txt", "idf_file.txt")

内容的提问来源于stack exchange,提问作者Andreas Rouvalis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 00:53:10