将Python字典转为Pandas DataFrame及gensim Word2Vec词向量处理问询
Hey Matt! Let's break down your two questions with clear, actionable examples so you can apply them right away.
The approach depends on how your dictionary is structured—here are the two most common scenarios:
Case 1: Dictionary maps column names to lists of values
If your dict has column headers as keys and lists of data as values, you can pass it directly to pd.DataFrame():
import pandas as pd # Example dictionary user_data = { 'username': ['1843', '866', '921'], 'join_date': ['2022-01-05', '2021-09-12', '2023-03-20'] } # Convert to DataFrame df = pd.DataFrame(user_data) print(df)
This will create a DataFrame where each key becomes a column, and the list values populate the rows.
Case 2: Dictionary maps row identifiers to row data
If your dict uses usernames/IDs as keys, and each value is a dict of column-value pairs, use from_dict() with orient='index':
user_rows = { '1843': {'vector': [0.082, -0.324, ...], 'account_type': 'premium'}, '866': {'vector': [-0.211, 0.105, ...], 'account_type': 'basic'} } df = pd.DataFrame.from_dict(user_rows, orient='index') # Move the row index (usernames) to a dedicated column df.reset_index(inplace=True) df.rename(columns={'index': 'username'}, inplace=True) print(df)
Based on your example, you want to turn your word vectors (tied to usernames/IDs) into a structured DataFrame. Here's how to do it efficiently:
Full Vocab Conversion
If you want to process every word/ID in your Word2Vec model's vocab:
import pandas as pd # Get all usernames/words from the model's vocab all_usernames = list(model.wv.vocab.keys()) # Extract the full vector for each username all_vectors = [model.wv[username] for username in all_usernames] # Create DataFrame: usernames as a column, vector components as separate columns df_vectors = pd.DataFrame(all_vectors, index=all_usernames) df_vectors.reset_index(inplace=True) df_vectors.rename(columns={'index': 'username'}, inplace=True) # Optional: Rename vector columns for clarity (e.g., vector_0, vector_1, ...) df_vectors.columns = ['username'] + [f'vector_{i}' for i in range(df_vectors.shape[1]-1)] print(df_vectors.head())
Sample-Specific Conversion (Like Your Example)
If you only need to process a subset of usernames (your sample list):
sample = ['1843', '866'] sample_vectors = [model.wv[w][:10] for w in sample] # Grab first 10 vector components # Build DataFrame with ID and vector columns df_sample = pd.DataFrame(sample_vectors, index=sample) df_sample.reset_index(inplace=True) df_sample.rename(columns={'index': 'ID'}, inplace=True) df_sample.columns = ['ID'] + [f'vector_{i}' for i in range(10)] print(df_sample)
This will output a DataFrame matching your example's structure, with the ID column and the first 10 vector values as separate columns.
内容的提问来源于stack exchange,提问作者Matt

