Python读取CSV时readline返回半数结果及IndexError问题排查
问题描述
我有一个包含6列的CSV文件,要做NLP处理,需要提取第6列(评论列)并转换成词的列表的列表。讲师提供了如下代码:
def read_twitter(fname): """ Read the given dataset into list and clean stop words. Args: fname (string): filename of Twitter Dataset Returns: list of list of words: we view each document as a list, including a list of all words """ twitter = [] with open(fname,encoding="utf-8") as f: for line in f: tweet = f.readline().split(",")[5] # YOUR CLEANING CODE HERE # - Clean tweet # - Split into list words # - Store list in twitter return twitter
调用代码:
twitter = read_twitter('twitter.csv')
未添加清洗代码时,本该返回空列表,却出现如下错误:
IndexError Traceback (most recent call last) in ~\AppData\Local\Temp\ipykernel_15784\2512851317.py in read_twitter(fname) 12 for line in f: 13 ---> 14 tweet = f.readline().split(",")[5] 15 16 IndexError: list index out of range.
修改代码后:
def read_twitter(fname): """ Read the given dataset into list and clean stop words. Args: fname (string): filename of Twitter Dataset Returns: list of list of words: we view each document as a list, including a list of all words """ twitter = [] with open(fname,encoding="utf-8") as f: for line in f: print(f.readline().split(",")[5]) return twitter twitter = read_twitter('twitter.csv')
能得到结果,但仅包含数据集的半数行,对readline()的行为和错误原因感到困惑。
问题分析与解决
1. 错误与半行结果的根源
for line in f会逐行迭代文件对象f,每次迭代时line已经指向当前行。此时再调用f.readline(),会直接读取下一行,导致循环跳过了一半的行(比如第1次循环读第1行,readline()读第2行;第2次循环读第3行,readline()读第4行,以此类推)。- 当文件行数为奇数时,最后一次循环里,
f.readline()会读到空字符串(文件末尾),空字符串用split(",")分割后得到[''],此时取索引5就会触发IndexError: list index out of range。
2. 正确实现方式
方法一:直接处理循环中的line变量
不需要额外调用readline(),直接用循环得到的line进行处理:
def read_twitter(fname): """ Read the given dataset into list and clean stop words. Args: fname (string): filename of Twitter Dataset Returns: list of list of words: we view each document as a list, including a list of all words """ twitter = [] # 示例停用词,可根据需求扩展 stop_words = {"the", "and", "is", "in", "a", "an"} with open(fname, encoding="utf-8") as f: # 跳过表头(如果CSV有表头的话) next(f) for line in f: # 分割行,取第6列(索引5) parts = line.strip().split(",") # 确保行有足够的列数,避免索引错误 if len(parts) >= 6: tweet = parts[5] # 清洗步骤:转小写、分词、过滤停用词 words = [word.lower() for word in tweet.split() if word.lower() not in stop_words] twitter.append(words) return twitter twitter = read_twitter('twitter.csv')
方法二:使用csv模块(更可靠)
如果CSV的评论列包含逗号(比如"I love, this movie"),直接用split(",")会分割错误,此时用Python内置的csv模块更安全:
import csv def read_twitter(fname): """ Read the given dataset into list and clean stop words. Args: fname (string): filename of Twitter Dataset Returns: list of list of words: we view each document as a list, including a list of all words """ twitter = [] stop_words = {"the", "and", "is", "in", "a", "an"} with open(fname, encoding="utf-8") as f: reader = csv.reader(f) # 跳过表头 next(reader) for row in reader: if len(row) >= 6: tweet = row[5] words = [word.lower() for word in tweet.split() if word.lower() not in stop_words] twitter.append(words) return twitter twitter = read_twitter('twitter.csv')
内容的提问来源于stack exchange,提问作者Alex
相关产品推荐
相关产品推荐

