You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复IndexError: index 1越界?文本分析代码问题排查

问题解答:列处理方式与条形图报错分析

一、列处理方式是否正确?

  • 你用index_col='Unnamed: 0'读取文件的做法没问题。如果原CSV里的Unnamed: 0列本来就是行索引(不是实际数据列),直接在读取时指定为索引列是最省事的方式,比事后用drop()/pop()处理更高效。
  • 之前drop()/pop()失败大概率是列名匹配问题:比如实际列名是Unnamed: 0,你可能写成了unnamed,大小写或后缀:0没对应上,导致操作没生效。

二、条形图报错IndexError: index 1 is out of bounds for axis 0 with size 1的原因

报错有两个关键问题:

  1. 错误拼接了整个DataFrame而非目标列:
    代码里" ".join(data)是对整个DataFrame做拼接,但DataFrame默认迭代的是列名,不是reviews列的内容。如果设置索引后你的DataFrame只剩reviews一列,那" ".join(data)的结果就是字符串"reviews",拆分后得到的Series只有一个元素,后面取freq_words[1]自然会触发索引越界。
  2. 条形图参数设置错误:
    freq_words是个pd.Series,它的索引是高频词,值是对应出现次数。plot.barh()不需要手动指定x和y,默认会用Series的索引当横轴标签,值当纵轴数值。你写的x=freq_words[0], y=freq_words[1]完全不符合Series的结构,进一步加重了错误。

三、修正后的代码

import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from wordcloud import WordCloud, STOPWORDS, ImageColorGenerator
import nltk
from nltk.sentiment.vader import SentimentIntensityAnalyzer
from nltk.corpus import stopwords
import string
import re
from nltk.tokenize import word_tokenize
from nltk.probability import FreqDist
from nltk.stem import WordNetLemmatizer
from nltk import ngrams
from collections import Counter
nltk.download('stopwords')
stemmer = nltk.SnowballStemmer('english')

# 读取文件,指定Unnamed:0为索引,方式正确
data = pd.read_csv('BA_reviews.csv', index_col='Unnamed: 0')
print(data.info())
print(data.head())

stopword = set(stopwords.words('english'))
def clean(text):
    text = str(text).lower()
    text = re.sub('\[.*?\]', '', text)
    text = re.sub('https?://\S+|www\.\S+', '', text)
    text = re.sub('<.*?>+', '', text)
    text = re.sub('[%s]' % re.escape(string.punctuation), '', text)
    text = re.sub('\n', '', text)
    text = re.sub('\w*\d\w*', '', text)
    text = [word for word in text.split(' ') if word not in stopword]
    text = " ".join(text)
    text = [stemmer.stem(word) for word in text.split(' ')]
    text = " ".join(text)
    text = re.sub('✅ Trip Verified |', '', text)
    text = re.sub('✅', '', text)
    text = re.sub('Trip Verified', '', text)
    text = re.sub('Verified', '', text)
    text = re.sub(' trip verifi', '', text)
    return text

data['reviews'] = data['reviews'].apply(clean)
print(data.head())
print(data.shape)

# 修正:只拼接reviews列的内容
freq_words = pd.Series(" ".join(data['reviews']).lower().split()).value_counts()[:50]
print(freq_words)

# 修正:直接调用barh,无需指定x/y参数
plt.figure(figsize=(10,10))
freq_words.plot.barh()
plt.title('Top 50 High Frequency Words')
plt.xlabel('Frequency')
plt.ylabel('Words')
plt.show()

内容的提问来源于stack exchange,提问作者Sumii

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 19:38:16