You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用W2V函数输出为DataFrame添加列及解决apply传参问题

问题描述

我有一个W2V函数,传入名字字符串会返回对应的向量数组,示例如下:

W2V('aamir')

返回结果:

array([ 0.12135 , -0.99132 ,  0.32347 ,  0.31334 ,  0.97446 , -0.67629 ,
        0.88606 , -0.11043 ,  0.79434 ,  1.4788  ,  0.53169 ,  0.95331 ,
       -1.1883  ,  0.82438 , -0.027177,  0.70081 ,  0.87467 , -0.095825,
       -0.5937  ,  1.4262  ,  0.2187  ,  1.1763  ,  1.6294  ,  0.91717 ,
       -0.086697,  0.16529 ,  0.19095 , -0.39362 , -0.40367 ,  0.83966 ,
       -0.25251 ,  0.46286 ,  0.82748 ,  0.93061 ,  1.136   ,  0.85616 ,
        0.34705 ,  0.65946 , -0.7143  ,  0.26379 ,  0.64717 ,  1.5633  ,
       -0.81238 , -0.44516 , -0.2979  ,  0.52601 , -0.41725 ,  0.086686,
        0.68263 , -0.15688 ], dtype=float32)

我有一个以Name为索引、仅包含Y列的DataFrame df1,结构如下:

Y
Name    
aamir   0
aaron   0
... ...
zulema  1
zuzana  1

我希望对每个Name值调用该函数生成对应特征列,目前用循环方法可实现但写法繁琐:

names = df1.index.to_list()

Lst = []
for name in names:
    Lst.append(W2V(name).tolist())
wv_df = pd.DataFrame(index=names, data=Lst)
wv_df.index.name = "Name"
wv_df.sort_index(inplace=True)

df1 = df1.merge(wv_df, how='inner', left_index=True, right_index=True)

我想改用apply()方法更高效地处理,但修改函数适配Series后得到错误结果,发现apply传入的是带索引标签的对象(如Name aamir)而非单纯的名字字符串aamir,请求解决该传参问题并提供高效实现方式。


解决方案

问题根源

你遇到的传参错误,是因为错误地对整个DataFrame的行(Series对象)调用了apply,此时传入函数的是包含索引标签和值的Series,而非纯名字字符串。正确的做法是直接针对索引值调用apply。

高效实现方式

方式1:apply配合索引转Series

将索引转为Series后调用apply,确保传入W2V的是纯名字字符串,同时自动生成特征列:

# 把索引转为Series,应用W2V并生成多列特征DataFrame
wv_features = df1.index.to_series().apply(lambda name: pd.Series(W2V(name)))
# 用join合并(比merge更简洁,因为索引一致)
df1 = df1.join(wv_features)

方式2:列表推导式+from_records(效率更优)

apply本质仍是逐元素循环,用列表推导式配合pd.DataFrame.from_records写法更简洁,效率和原循环一致:

# 直接生成特征数组列表,转为DataFrame并复用原索引
wv_features = pd.DataFrame.from_records(
    [W2V(name) for name in df1.index],
    index=df1.index
)
# 合并到原DataFrame
df1 = df1.join(wv_features)

补充优化

如果需要给特征列设置有意义的名称,可以在生成DataFrame时指定columns参数,比如:

# 假设向量长度为50,生成wv_0到wv_49的列名
wv_features = pd.DataFrame.from_records(
    [W2V(name) for name in df1.index],
    index=df1.index,
    columns=[f'wv_{i}' for i in range(50)]
)

内容的提问来源于stack exchange,提问作者Brian Feeny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 14:35:35