You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中列表与整数除法报错及词向量求和平均问题排查

Fixing Word2Vec Vector Sum & Average Calculation Error

Let's break down what's going wrong and fix your code properly:

Why You're Getting the Error

The core issue is how you're handling vector storage and arithmetic:

  • You initialized total as an empty list, but model[kata] returns a numpy ndarray (the word vector). Adding a numpy array to a list doesn't do element-wise vector addition—it causes type mismatches, and eventually trying to divide a list by an integer throws that TypeError.
  • Initializing total=[100] is also incorrect because it's a 1-element list, which doesn't match the 100-dimensional vectors from your model. This leads to invalid vector operations and wrong results.

Correct Approach

We need to use numpy arrays for all vector operations, since Word2Vec vectors are already numpy arrays. Here's how to adjust your code:

  1. Import numpy to handle array operations.
  2. Initialize total as a zero-filled numpy array matching your vector dimension (100 in your case).
  3. Use element-wise addition (+=) to accumulate vectors correctly.
  4. Add safeguards to avoid division by zero (in case none of the words are found in the model).

Modified Working Code

import re
import csv
import gensim
import pandas as pd
import numpy as np  # Add this import

def processing(kata):
    words = re.sub(r'([^\s\w]|_)', '', str(kata))
    words = re.sub(r'[0-9]+', '', words)
    return words

path = './model_terbaru/idwiki_word2vec_100.model'
model = gensim.models.word2vec.Word2Vec.load(path)
data = pd.read_csv('data/coba_data1.csv', encoding="ISO-8859-1")

for index, row in data.iterrows():
    # Initialize total as 100-dimensional zero array (matches your model's vector size)
    total = np.zeros(100)
    kalimat = row[0]
    valid_word_count = 0  # Track how many words were actually found in the model
    
    # Preprocess the sentence
    words = re.sub(r'([^\s\w]|_)', '', str(kalimat))
    words = re.sub(r'[0-9]+', '', words)
    
    for word in words.split():
        kata = word.lower()
        try:
            vector = model[kata]  # No need for extra quotes here
            total += vector  # Element-wise vector addition
            valid_word_count += 1
            print(f"Vector for {kata}: {vector}")
            print(f"Running total: {total}")
        except KeyError:  # Catch specific KeyError instead of bare except
            print(f"Word '{kata}' not found in model, skipping")
            pass
    
    print(f"Total valid words found: {valid_word_count}")
    
    # Calculate average only if we have valid words
    if valid_word_count > 0:
        rata = total / valid_word_count  # Element-wise division
        print(f"Average vector: {rata}")
        
        # Optional: Write to CSV
        # with open('data/vector_training.csv', 'a', newline='') as ok:
        #     writer = csv.writer(ok)
        #     writer.writerow(rata.tolist())  # Convert numpy array to list for CSV
    else:
        print("No valid words found in this sentence, skipping average calculation")

Key Changes Explained

  • np.zeros(100): Initializes total as a 100-dimensional zero array, matching the output of your Word2Vec model.
  • total += vector: Properly accumulates each word's vector using element-wise addition (the correct way to sum vectors).
  • valid_word_count: Tracks how many words were actually found in the model (instead of using len(words.split()), which counts all words including those not in the model). This gives a more accurate average.
  • Specific KeyError catch: Avoids catching unexpected errors with a bare except.
  • Division safeguard: Prevents division by zero when no words are found in the model.

内容的提问来源于stack exchange,提问作者Muhammad Rusli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:59:16