Pandas新增合并列报错:ValueError: 值长度与索引长度不匹配
问题
尝试在CSV数据集新增一列,将title、authors、publisher三列内容合并到新列中,编写了combine_features函数实现逻辑,但运行时报错:
ValueError: Length of values (1) does not match length of index (11123)
已尝试df.reset_index(inplace=True,drop=True),问题依旧。
代码如下:
import numpy as np import pandas as pd df = pd.read_csv('books.csv', encoding='unicode_escape', error_bad_lines=False) #List of columns to keep columns =['title', 'authors', 'publisher'] #Function to combine the columns/features def combine_features(data): features = [] for i in range(0, data.shape[0]): features.append( data['title'][i] +' '+data['authors'][i]+' '+data['publisher'][i]) return features #Column to store the combined features df['combined_features'] =combine_features(df) #Show data df
完整报错信息:
ValueError Traceback (most recent call last) <ipython-input-24-40cc76d3cd85> in <module> 1 #Create a column to store the combined features ----> 2 df['combined_features'] =combine_features(df) 3 df 3 frames /usr/local/lib/python3.8/dist-packages/pandas/core/frame.py in __setitem__(self, key, value) 3610 else: 3611 # set column -> 3612 self._set_item(key, value) 3613 3614 def _setitem_slice(self, key: slice, value): /usr/local/lib/python3.8/dist-packages/pandas/core/frame.py in _set_item(self, key, value) 3782 ensure homogeneity. 3783 """ -> 3784 value = self._sanitize_column(value) 3785 3786 if ( /usr/local/lib/python3.8/dist-packages/pandas/core/frame.py in _sanitize_column(self, value) 4507 4508 if is_list_like(value): -> 4509 com.require_length_match(value, self.index) 4510 return sanitize_array(value, self.index, copy=True, allow_2d=True) 4511 /usr/local/lib/python3.8/dist-packages/pandas/core/common.py in require_length_match(data, index) 529 """ 530 if len(data) != len(index): -> 531 raise ValueError( 532 "Length of values " 533 f"({len(data)}) " ValueError: Length of values (1) does not match length of index (11123)
解决方法
错误原因
combine_features函数里的return features缩进错误,它被放在了for循环内部。这导致循环仅执行一次(添加第一行的合并内容)就直接返回列表,最终features只有1个元素,和DataFrame的11123行长度不匹配,触发报错。
修正后的代码
把return features移到for循环外面,确保遍历完所有行再返回完整列表:
import numpy as np import pandas as pd df = pd.read_csv('books.csv', encoding='unicode_escape', error_bad_lines=False) #Function to combine the columns/features def combine_features(data): features = [] for i in range(0, data.shape[0]): features.append( data['title'][i] +' '+data['authors'][i]+' '+data['publisher'][i]) # 将return移至循环外部 return features #Column to store the combined features df['combined_features'] = combine_features(df) #Show data df
更高效的写法(推荐)
用Pandas的向量操作替代循环,代码更简洁且运行效率更高:
import pandas as pd df = pd.read_csv('books.csv', encoding='unicode_escape', error_bad_lines=False) # 直接通过向量拼接合并三列 df['combined_features'] = df['title'] + ' ' + df['authors'] + ' ' + df['publisher'] # 查看结果 df
内容的提问来源于stack exchange,提问作者mudgey
相关产品推荐
相关产品推荐

