You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中DataFrame字符串列转Word2Vec可用列表格式的方法

Solution to Convert String Lists to Word2Vec-Ready Format

Got it, let's fix this so your Word2Vec code runs smoothly. The core issue here is that your df['set'] column contains string representations of lists, not actual lists of strings—which is what Gensim's Word2Vec expects (it needs sequences of string "tokens").

Step-by-Step Fix

We'll create a simple parsing function to convert each string into a list of string elements, then apply it to your DataFrame column.

  1. Define the parsing function
    This function handles both comma-separated formats (with or without spaces) by stripping brackets, splitting elements, cleaning whitespace, and converting each item to a string:

    def parse_string_list(s):
        # Remove the surrounding brackets
        cleaned = s.strip('[]')
        # Split into individual elements (handles commas with/without spaces)
        elements = cleaned.split(',')
        # Strip whitespace from each element and convert to string
        return [elem.strip() for elem in elements]
    
  2. Apply the function to your DataFrame column
    Use pandas' apply() method to transform every entry in df['set']:

    df['set'] = df['set'].apply(parse_string_list)
    

Full Working Example

Here's a complete code snippet including sample data, conversion, and Word2Vec initialization:

import pandas as pd
from gensim.models import Word2Vec

# Sample DataFrame matching your input format
data = {
    'set': [
        "[911,3040]",
        "[130055, 99832, 62131]",
        "[19397, 3987, 5330, 14781]",
        "[76514, 70178, 70301, 76545]",
        "[79185, 38367, 131155, 79433]"
    ]
}
df = pd.DataFrame(data)

# Parse the string lists
def parse_string_list(s):
    cleaned = s.strip('[]')
    elements = cleaned.split(',')
    return [elem.strip() for elem in elements]

df['set'] = df['set'].apply(parse_string_list)

# Now your Word2Vec code will work!
# Note: In newer Gensim versions, use `vector_size` instead of `size`
model = Word2Vec(df['set'], vector_size=100)

Verification

After running the conversion, checking df['set'][0] will return ['911', '3040']—exactly the format you need for Word2Vec to process the tokens correctly.

内容的提问来源于stack exchange,提问作者Adi Milrad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:56:34