Python 2.7读取指定路径多文本文件,TF-IDF前十高频词报错求助
修复Python 2.7中TF-IDF提取文本Top10高频词的问题
Hey there! 作为Python 2.7新手折腾TF-IDF确实容易踩坑,我来帮你梳理常见问题和修复方案~毕竟Python 2.7在编码、第三方库兼容上有不少特殊点,这些往往是出错的根源。
常见出错原因&对应修复要点
- 路径遍历遗漏文件:没正确递归遍历目标路径下所有
.txt文件,或者判断文件后缀的逻辑有误 - 编码爆炸:Python 2.7默认用ASCII编码处理字符串,读取含中文/特殊字符的文本时很容易触发解码错误
- sklearn版本不兼容:Python 2.7只能支持
scikit-learn 0.20.4及以下版本,高版本直接装不上会报错 - TF-IDF逻辑混淆:比如误把全局TF-IDF当成单个文件的词频,或者提取Top10时排序方向搞反
修复后的完整代码(带注释)
import os from sklearn.feature_extraction.text import TfidfVectorizer import numpy as np # 替换成你的目标文件路径 TARGET_DIR = "/path/to/your/text/files" def load_text_files(dir_path): text_contents = [] file_names = [] # 递归遍历所有子目录下的.txt文件 for root, _, files in os.walk(dir_path): for file in files: if file.lower().endswith(".txt"): # 小写后缀避免大小写问题 full_path = os.path.join(root, file) file_names.append(file) # 处理编码问题,根据你的文件实际编码调整(比如gbk) try: with open(full_path, "r") as f: # Python2.7需把bytes转成Unicode才能被TF-IDF处理 content = f.read().decode("utf-8") text_contents.append(content) except UnicodeDecodeError: print(f"⚠️ 文件 {full_path} 编码异常,已跳过") continue except IOError: print(f"⚠️ 无法读取文件 {full_path},已跳过") continue return text_contents, file_names def extract_top10_tfidf(texts, filenames): # 初始化TF-IDF向量器,处理英文用默认停用词;中文需替换成中文停用词表 tfidf_vectorizer = TfidfVectorizer(stop_words="english") tfidf_matrix = tfidf_vectorizer.fit_transform(texts) # 获取所有特征词(即所有文件里的词汇) feature_words = np.array(tfidf_vectorizer.get_feature_names()) # 逐个文件提取Top10高频词 for idx, name in enumerate(filenames): # 获取当前文件的TF-IDF分数数组 scores = tfidf_matrix[idx].toarray().flatten() # 按分数降序取前10个词的索引 top_indices = scores.argsort()[-10:][::-1] top_words = feature_words[top_indices] top_scores = scores[top_indices] print(f"\n📄 文件 {name} 的Top10 TF-IDF高频词:") for word, score in zip(top_words, top_scores): print(f" {word}: {score:.4f}") if __name__ == "__main__": texts, filenames = load_text_files(TARGET_DIR) if texts: extract_top10_tfidf(texts, filenames) else: print("❌ 未找到任何可正常读取的文本文件")
额外注意事项
- 安装兼容的sklearn版本:执行以下命令安装适配Python2.7的版本
pip install scikit-learn==0.20.4 - 中文文本处理:如果你的文件是中文,需要替换
stop_words为中文停用词表(比如从网上下载中文停用词文本,读取后转成列表传入TfidfVectorizer) - 错误排查:如果运行还是报错,把具体的错误信息(比如
ImportError、UnicodeError)贴出来,我可以帮你更精准地定位问题~
内容的提问来源于stack exchange,提问作者Leonin Joy Pantuan
相关产品推荐
相关产品推荐

