如何用W2V函数输出为DataFrame添加列及解决apply传参问题
问题描述
我有一个W2V函数,传入名字字符串会返回对应的向量数组,示例如下:
W2V('aamir')
返回结果:
array([ 0.12135 , -0.99132 , 0.32347 , 0.31334 , 0.97446 , -0.67629 , 0.88606 , -0.11043 , 0.79434 , 1.4788 , 0.53169 , 0.95331 , -1.1883 , 0.82438 , -0.027177, 0.70081 , 0.87467 , -0.095825, -0.5937 , 1.4262 , 0.2187 , 1.1763 , 1.6294 , 0.91717 , -0.086697, 0.16529 , 0.19095 , -0.39362 , -0.40367 , 0.83966 , -0.25251 , 0.46286 , 0.82748 , 0.93061 , 1.136 , 0.85616 , 0.34705 , 0.65946 , -0.7143 , 0.26379 , 0.64717 , 1.5633 , -0.81238 , -0.44516 , -0.2979 , 0.52601 , -0.41725 , 0.086686, 0.68263 , -0.15688 ], dtype=float32)
我有一个以Name为索引、仅包含Y列的DataFrame df1,结构如下:
Y Name aamir 0 aaron 0 ... ... zulema 1 zuzana 1
我希望对每个Name值调用该函数生成对应特征列,目前用循环方法可实现但写法繁琐:
names = df1.index.to_list() Lst = [] for name in names: Lst.append(W2V(name).tolist()) wv_df = pd.DataFrame(index=names, data=Lst) wv_df.index.name = "Name" wv_df.sort_index(inplace=True) df1 = df1.merge(wv_df, how='inner', left_index=True, right_index=True)
我想改用apply()方法更高效地处理,但修改函数适配Series后得到错误结果,发现apply传入的是带索引标签的对象(如Name aamir)而非单纯的名字字符串aamir,请求解决该传参问题并提供高效实现方式。
解决方案
问题根源
你遇到的传参错误,是因为错误地对整个DataFrame的行(Series对象)调用了apply,此时传入函数的是包含索引标签和值的Series,而非纯名字字符串。正确的做法是直接针对索引值调用apply。
高效实现方式
方式1:apply配合索引转Series
将索引转为Series后调用apply,确保传入W2V的是纯名字字符串,同时自动生成特征列:
# 把索引转为Series,应用W2V并生成多列特征DataFrame wv_features = df1.index.to_series().apply(lambda name: pd.Series(W2V(name))) # 用join合并(比merge更简洁,因为索引一致) df1 = df1.join(wv_features)
方式2:列表推导式+from_records(效率更优)
apply本质仍是逐元素循环,用列表推导式配合pd.DataFrame.from_records写法更简洁,效率和原循环一致:
# 直接生成特征数组列表,转为DataFrame并复用原索引 wv_features = pd.DataFrame.from_records( [W2V(name) for name in df1.index], index=df1.index ) # 合并到原DataFrame df1 = df1.join(wv_features)
补充优化
如果需要给特征列设置有意义的名称,可以在生成DataFrame时指定columns参数,比如:
# 假设向量长度为50,生成wv_0到wv_49的列名 wv_features = pd.DataFrame.from_records( [W2V(name) for name in df1.index], index=df1.index, columns=[f'wv_{i}' for i in range(50)] )
内容的提问来源于stack exchange,提问作者Brian Feeny
相关产品推荐
相关产品推荐

