Python批量Tokenize目录多文件异常:循环仅处理特定文本
问题分析与修复方案
你的代码存在几个关键问题导致无法批量处理文件:
- 未导入
os模块,调用os.listdir会直接报错 - 文件路径拼接错误:
open(filename)只会在当前工作目录查找文件,需要拼接words_input_dir和filename得到完整路径 word_tokensize是未定义的函数,属于多余或笔误代码tokenizer.encode_plus写在循环外部,且尝试读取已关闭的文件句柄(with块结束后文件已自动关闭)- 没有对每个文件的内容执行Tokenize并输出/保存结果
修正后的代码
from transformers import AutoTokenizer, AutoModelForSequenceClassification import torch import os # 新增导入os模块 # 初始化tokenizer和模型(仅做tokenize可省略模型加载,这里保留原有代码) tokenizer = AutoTokenizer.from_pretrained("joeddav/distilbert-base-uncased-go-emotions-student") model = AutoModelForSequenceClassification.from_pretrained("joeddav/distilbert-base-uncased-go-emotions-student") words_input_dir = "/content/sample_data/" # 遍历目录下的所有txt文件 for filename in os.listdir(words_input_dir): if filename.endswith(".txt"): # 拼接完整文件路径 file_path = os.path.join(words_input_dir, filename) with open(file_path, "r", encoding="utf-8") as input_file: text = input_file.read() # 对当前文件内容执行tokenize tokens = tokenizer.encode_plus( text, add_special_tokens=False, return_tensors='pt' ) # 输出当前文件的token数量 print(f"文件 {filename} 的token数量: {len(tokens['input_ids'][0])}") # 可选:保存token结果到文件 # with open(f"{filename}_tokens.txt", "w") as f: # f.write(str(tokens['input_ids'][0].tolist()))
关键修改说明
- 新增
import os,解决文件遍历的依赖问题 - 使用
os.path.join拼接完整文件路径,确保能正确定位目标目录下的文件 - 移除未定义的
word_tokensize调用,直接读取文件内容后执行tokenize - 将
tokenizer.encode_plus放入with块内部,确保文件句柄有效,且每个文件都能被单独处理 - 针对每个文件输出处理结果,替换原有的单次打印逻辑
- 可选添加结果保存逻辑,方便后续查看处理后的token数据
内容的提问来源于stack exchange,提问作者Seb
相关产品推荐
相关产品推荐

