You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将Python字典转为Pandas DataFrame及gensim Word2Vec词向量处理问询

Hey Matt! Let's break down your two questions with clear, actionable examples so you can apply them right away.

1. Converting a Python Dictionary to a Pandas DataFrame

The approach depends on how your dictionary is structured—here are the two most common scenarios:

Case 1: Dictionary maps column names to lists of values

If your dict has column headers as keys and lists of data as values, you can pass it directly to pd.DataFrame():

import pandas as pd

# Example dictionary
user_data = {
    'username': ['1843', '866', '921'],
    'join_date': ['2022-01-05', '2021-09-12', '2023-03-20']
}

# Convert to DataFrame
df = pd.DataFrame(user_data)
print(df)

This will create a DataFrame where each key becomes a column, and the list values populate the rows.

Case 2: Dictionary maps row identifiers to row data

If your dict uses usernames/IDs as keys, and each value is a dict of column-value pairs, use from_dict() with orient='index':

user_rows = {
    '1843': {'vector': [0.082, -0.324, ...], 'account_type': 'premium'},
    '866': {'vector': [-0.211, 0.105, ...], 'account_type': 'basic'}
}

df = pd.DataFrame.from_dict(user_rows, orient='index')
# Move the row index (usernames) to a dedicated column
df.reset_index(inplace=True)
df.rename(columns={'index': 'username'}, inplace=True)
print(df)
2. Converting Gensim Word2Vec Vectors to a Pandas DataFrame

Based on your example, you want to turn your word vectors (tied to usernames/IDs) into a structured DataFrame. Here's how to do it efficiently:

Full Vocab Conversion

If you want to process every word/ID in your Word2Vec model's vocab:

import pandas as pd

# Get all usernames/words from the model's vocab
all_usernames = list(model.wv.vocab.keys())
# Extract the full vector for each username
all_vectors = [model.wv[username] for username in all_usernames]

# Create DataFrame: usernames as a column, vector components as separate columns
df_vectors = pd.DataFrame(all_vectors, index=all_usernames)
df_vectors.reset_index(inplace=True)
df_vectors.rename(columns={'index': 'username'}, inplace=True)

# Optional: Rename vector columns for clarity (e.g., vector_0, vector_1, ...)
df_vectors.columns = ['username'] + [f'vector_{i}' for i in range(df_vectors.shape[1]-1)]
print(df_vectors.head())

Sample-Specific Conversion (Like Your Example)

If you only need to process a subset of usernames (your sample list):

sample = ['1843', '866']
sample_vectors = [model.wv[w][:10] for w in sample]  # Grab first 10 vector components

# Build DataFrame with ID and vector columns
df_sample = pd.DataFrame(sample_vectors, index=sample)
df_sample.reset_index(inplace=True)
df_sample.rename(columns={'index': 'ID'}, inplace=True)
df_sample.columns = ['ID'] + [f'vector_{i}' for i in range(10)]

print(df_sample)

This will output a DataFrame matching your example's structure, with the ID column and the first 10 vector values as separate columns.

内容的提问来源于stack exchange,提问作者Matt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:08:17