Single Positional Indexer越界:对话AI中Question转Regresponse报错
对话AI索引越界错误排查与修复
问题描述
我正在开发一款对话AI,作为Python新手,代码和调试方法大多来自网络。目前除了Question转Regresponse的场景外,其他功能都正常运行,但反复出现Single Positional Indexer out-of-bounds错误,尝试多种调试方式均无效。
代码实现
import pandas as pd from sklearn.feature_extraction.text import CountVectorizer from sklearn.naive_bayes import MultinomialNB import random # Load the data from the CSV file data = pd.read_csv("conversational_english.csv") # Split the data into training and testing sets training_data = data[:int(0.8 * len(data))] testing_data = data[int(0.8 * len(data)):] # Convert the text into numerical feature vectors using CountVectorizer vectorizer = CountVectorizer() text_features = vectorizer.fit_transform(training_data['text']) # Train a Naive Bayes classifier on the training data classifier = MultinomialNB().fit(text_features, training_data['label']) # Continuously get user input and generate a response based on the label while True: user_input = input("Enter a conversational text: ") if user_input == "exit": break user_input_features = vectorizer.transform([user_input]) predicted_label = classifier.predict(user_input_features)[0] if predicted_label == 'feeling_question': predicted_label = 'feeling_response' if predicted_label == 'question': predicted_label = 'regresponse' if predicted_label == 'joke_request': predicted_label = 'joke' print(predicted_label) response = data.loc[data['label'] == predicted_label, 'text'].iloc[0] print("Response:", response) correct = input("Is this response correct? (yes/no): ") if correct == "no": new_label = input("Enter the correct label for the user input: ") if user_input in data['text'].values: continue else: new_data = pd.DataFrame({'text': [user_input], 'label': [new_label]}) data = data.append(new_data, ignore_index=True) data.to_csv("conversational_english.csv", index=False) else: if user_input in data['text'].values: continue else: new_data = pd.DataFrame({'text': [user_input], 'label': [predicted_label]}) data = data.append(new_data, ignore_index=True) data.to_csv("conversational_english.csv", index=False)
CSV数据
text,label, Can you tell me a joke?,joke_request, I'm not sure because I am an AI, regresponse, Hi how are you doing today?,greeting, Why did the chicken cross the road?,joke, Where is the bus stop?,question, how are you?,feeling_question, What are you doing?,feeling_question, I am an AI. How should I know?, regresponse, How are you?,feeling_question, Goodbye,farewell, Bye,farewell, "Two whales walk into a bar. One says ""oOooOoooOOh"". The other says ""What the hell Jim""",joke, Hi!,greeting, I'm not sure. I am an AI, regresponse, Hello!,greeting, Greetings!,greeting, what's up?,question, I'm doing good!,feeling_response, What's up?,question, Tell me a joke,joke_request, Tell me something funny,joke_request, I'm doing great!,feeling_reponse, I am an AI. I don't know., regresponse, Not bad!,feeling_response, Hi,greeting, Hello,greeting, how's it going?,question, My name is Zach,name, how's it goin,question, what are you up to?,question, I don't know I am and AI, regresponse, How are you,feeling_question, How's it goin,question, How do you do,feeling_question, "A horse walks into a bar and the bartender says ""why the long face""",joke, goodbye,farewell, hello,greeting, Good evening,greeting, I'm not sure., regresponse, hello!,greeting, What's up,question, Heyo,greeting, hi,greeting, I am an AI. I can't help you., regresponse, How's it going?,question, I don't know because I am an AI., regresponse, What's up!,question,
报错信息
Traceback (most recent call last): File "C:\Users\Kyn\OneDrive\Documents\test-rep\Python AI\response.py", line 35, in <module> response = data.loc[data['label'] == predicted_label, 'text'].iloc[0] File "C:\Users\Kyn\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\pandas\core\indexing.py", line 1073, in __getitem__ return self._getitem_axis(maybe_callable, axis=axis) File "C:\Users\Kyn\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\pandas\core\indexing.py", line 1625, in _getitem_axis self._validate_integer(key, axis) File "C:\Users\Kyn\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\pandas\core\indexing.py", line 1557, in _validate_integer raise IndexError("single positional indexer is out-of-bounds")
错误原因
核心问题出在CSV数据的标签格式不统一:
- 所有
regresponse标签前面都有一个多余的空格(比如regresponse),但代码中转换后的标签是regresponse(无空格),导致data.loc[data['label'] == 'regresponse']找不到任何匹配行,空的Series调用.iloc[0]就会触发索引越界错误。 - CSV中存在拼写错误:
feeling_reponse应为feeling_response,这会导致后续匹配该标签时同样可能出现无结果的情况。
修复方案
1. 清理CSV标签数据
打开conversational_english.csv,做两处修改:
- 把所有
regresponse(带前置空格)替换为regresponse(无空格) - 把
feeling_reponse替换为feeling_response
2. 代码层面增加容错与优化
(1)加载数据时自动清理标签
在读取CSV后,添加一行代码自动去除标签的前后空格,避免后续再出现空格问题:
data = pd.read_csv("conversational_english.csv") # 清理label列的前后空白字符 data['label'] = data['label'].str.strip()
(2)增加空结果判断,避免索引越界
替换原来获取响应的代码,先检查是否有匹配行,没有则返回默认提示,同时随机选择响应让对话更自然:
# 替换原response = ...那一行 matching_responses = data.loc[data['label'] == predicted_label, 'text'] if not matching_responses.empty: # 随机选一个匹配的响应 response = random.choice(matching_responses.tolist()) else: response = "Sorry, I don't have a suitable response right now."
(3)简化标签转换逻辑
用字典替代多个if判断,代码更简洁易维护:
# 替换原三个if标签转换的代码 label_mapping = { 'feeling_question': 'feeling_response', 'question': 'regresponse', 'joke_request': 'joke' } predicted_label = label_mapping.get(predicted_label, predicted_label)
内容的提问来源于stack exchange,提问作者Zach
相关产品推荐
相关产品推荐

