Python中DataFrame字符串列转Word2Vec可用列表格式的方法
Got it, let's fix this so your Word2Vec code runs smoothly. The core issue here is that your df['set'] column contains string representations of lists, not actual lists of strings—which is what Gensim's Word2Vec expects (it needs sequences of string "tokens").
Step-by-Step Fix
We'll create a simple parsing function to convert each string into a list of string elements, then apply it to your DataFrame column.
Define the parsing function
This function handles both comma-separated formats (with or without spaces) by stripping brackets, splitting elements, cleaning whitespace, and converting each item to a string:def parse_string_list(s): # Remove the surrounding brackets cleaned = s.strip('[]') # Split into individual elements (handles commas with/without spaces) elements = cleaned.split(',') # Strip whitespace from each element and convert to string return [elem.strip() for elem in elements]Apply the function to your DataFrame column
Use pandas'apply()method to transform every entry indf['set']:df['set'] = df['set'].apply(parse_string_list)
Full Working Example
Here's a complete code snippet including sample data, conversion, and Word2Vec initialization:
import pandas as pd from gensim.models import Word2Vec # Sample DataFrame matching your input format data = { 'set': [ "[911,3040]", "[130055, 99832, 62131]", "[19397, 3987, 5330, 14781]", "[76514, 70178, 70301, 76545]", "[79185, 38367, 131155, 79433]" ] } df = pd.DataFrame(data) # Parse the string lists def parse_string_list(s): cleaned = s.strip('[]') elements = cleaned.split(',') return [elem.strip() for elem in elements] df['set'] = df['set'].apply(parse_string_list) # Now your Word2Vec code will work! # Note: In newer Gensim versions, use `vector_size` instead of `size` model = Word2Vec(df['set'], vector_size=100)
Verification
After running the conversion, checking df['set'][0] will return ['911', '3040']—exactly the format you need for Word2Vec to process the tokens correctly.
内容的提问来源于stack exchange,提问作者Adi Milrad

