You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从多个pandas DataFrame提取列创建新表时遇ValueError错误

解决pandas合并唯一值列时的长度匹配错误

我之前在Python2.7的旧版pandas里也碰到过一模一样的问题——明明反复确认三个数组长度完全一致,但赋值时还是跳出ValueError: Length of values does not match length of index的报错。其实核心原因是numpy数组和Python列表在DataFrame赋值时的索引处理逻辑差异,再加上旧版pandas的一些小bug导致的。

问题拆解

你代码里的sorted(df1.columndf1.unique())返回的是Python原生列表,而df2.columndf2.unique()这类调用返回的是numpy.ndarray对象。在旧版pandas中,当你先给空DataFrame赋值一个列表(此时DataFrame会生成0-based的连续索引),再赋值numpy数组时,pandas会尝试对齐数组的“隐式索引”——虽然数组长度和DataFrame索引长度一致,但旧版pandas可能误判数组的索引结构,导致对齐失败。哪怕你转成Series,如果不手动指定索引,也可能因为默认索引的对齐逻辑出问题。

两种可行的解决方案

方案1:统一转成Python列表赋值

把所有unique()返回的numpy数组都转成列表,彻底规避索引对齐的麻烦:

import pandas as pd

# 初始化空DataFrame
newdf = pd.DataFrame()

# 将unique结果转成列表后再处理赋值
newdf['col1'] = sorted(df1.columndf1.unique().tolist())
newdf['col2'] = df2.columndf2.unique().tolist()
newdf['col3'] = df3.columndf3.unique().tolist()

方案2:用pd.concat合并带相同索引的Series

手动指定所有Series的索引和第一个列保持一致,确保合并时不会出现索引错位:

import pandas as pd

# 生成第一个列的Series,作为基准索引
s_col1 = pd.Series(sorted(df1.columndf1.unique()), name='col1')

# 给另外两个Series指定和s_col1完全相同的索引
s_col2 = pd.Series(df2.columndf2.unique(), name='col2', index=s_col1.index)
s_col3 = pd.Series(df3.columndf3.unique(), name='col3', index=s_col1.index)

# 横向合并三个Series得到新DataFrame
newdf = pd.concat([s_col1, s_col2, s_col3], axis=1)

额外提醒

Python2.7和对应旧版pandas的兼容性问题确实不少,如果条件允许,建议升级到Python3.x和较新版本的pandas;但如果必须在现有环境下工作,上面两种方法都能稳定解决你的问题。

内容的提问来源于stack exchange,提问作者sato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:22:41