You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取CSV时readline返回半数结果及IndexError问题排查

问题描述

我有一个包含6列的CSV文件,要做NLP处理,需要提取第6列(评论列)并转换成词的列表的列表。讲师提供了如下代码:

def read_twitter(fname):
    """ Read the given dataset into list and clean stop words. 
    
    Args: 
        fname (string): filename of Twitter Dataset
        
    Returns:
        list of list of words: we view each document as a list, including a list of all words 
    """
    twitter = []
    with open(fname,encoding="utf-8") as f:
        for line in f:
            tweet = f.readline().split(",")[5]
            
            # YOUR CLEANING CODE HERE
            #    - Clean tweet
            #    - Split into list words
            #    - Store list in twitter
            
    return twitter

调用代码:

twitter = read_twitter('twitter.csv')

未添加清洗代码时,本该返回空列表,却出现如下错误:

IndexError                                Traceback (most recent call last)
 in 

~\AppData\Local\Temp\ipykernel_15784\2512851317.py in read_twitter(fname)

 12         for line in f:

 13 

---> 14             tweet = f.readline().split(",")[5]

 15 

 16 

IndexError: list index out of range.

修改代码后:

def read_twitter(fname):
    """ Read the given dataset into list and clean stop words. 
    
    Args: 
        fname (string): filename of Twitter Dataset
        
    Returns:
        list of list of words: we view each document as a list, including a list of all words 
    """
    twitter = []
    with open(fname,encoding="utf-8") as f:
        for line in f:
            print(f.readline().split(",")[5])
            
    return twitter
twitter = read_twitter('twitter.csv')

能得到结果,但仅包含数据集的半数行,对readline()的行为和错误原因感到困惑。

问题分析与解决

1. 错误与半行结果的根源

  • for line in f会逐行迭代文件对象f,每次迭代时line已经指向当前行。此时再调用f.readline(),会直接读取下一行,导致循环跳过了一半的行(比如第1次循环读第1行,readline()读第2行;第2次循环读第3行,readline()读第4行,以此类推)。
  • 当文件行数为奇数时,最后一次循环里,f.readline()会读到空字符串(文件末尾),空字符串用split(",")分割后得到[''],此时取索引5就会触发IndexError: list index out of range。

2. 正确实现方式

方法一:直接处理循环中的line变量

不需要额外调用readline(),直接用循环得到的line进行处理:

def read_twitter(fname):
    """ Read the given dataset into list and clean stop words. 
    
    Args: 
        fname (string): filename of Twitter Dataset
        
    Returns:
        list of list of words: we view each document as a list, including a list of all words 
    """
    twitter = []
    # 示例停用词,可根据需求扩展
    stop_words = {"the", "and", "is", "in", "a", "an"}
    
    with open(fname, encoding="utf-8") as f:
        # 跳过表头(如果CSV有表头的话)
        next(f)
        for line in f:
            # 分割行,取第6列(索引5)
            parts = line.strip().split(",")
            # 确保行有足够的列数,避免索引错误
            if len(parts) >= 6:
                tweet = parts[5]
                # 清洗步骤:转小写、分词、过滤停用词
                words = [word.lower() for word in tweet.split() if word.lower() not in stop_words]
                twitter.append(words)
    return twitter

twitter = read_twitter('twitter.csv')

方法二:使用csv模块(更可靠)

如果CSV的评论列包含逗号(比如"I love, this movie"),直接用split(",")会分割错误,此时用Python内置的csv模块更安全:

import csv

def read_twitter(fname):
    """ Read the given dataset into list and clean stop words. 
    
    Args: 
        fname (string): filename of Twitter Dataset
        
    Returns:
        list of list of words: we view each document as a list, including a list of all words 
    """
    twitter = []
    stop_words = {"the", "and", "is", "in", "a", "an"}
    
    with open(fname, encoding="utf-8") as f:
        reader = csv.reader(f)
        # 跳过表头
        next(reader)
        for row in reader:
            if len(row) >= 6:
                tweet = row[5]
                words = [word.lower() for word in tweet.split() if word.lower() not in stop_words]
                twitter.append(words)
    return twitter

twitter = read_twitter('twitter.csv')

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 04:27:20