Python中列表与整数除法报错及词向量求和平均问题排查
Fixing Word2Vec Vector Sum & Average Calculation Error
Let's break down what's going wrong and fix your code properly:
Why You're Getting the Error
The core issue is how you're handling vector storage and arithmetic:
- You initialized
totalas an emptylist, butmodel[kata]returns a numpy ndarray (the word vector). Adding a numpy array to a list doesn't do element-wise vector addition—it causes type mismatches, and eventually trying to divide a list by an integer throws thatTypeError. - Initializing
total=[100]is also incorrect because it's a 1-element list, which doesn't match the 100-dimensional vectors from your model. This leads to invalid vector operations and wrong results.
Correct Approach
We need to use numpy arrays for all vector operations, since Word2Vec vectors are already numpy arrays. Here's how to adjust your code:
- Import numpy to handle array operations.
- Initialize
totalas a zero-filled numpy array matching your vector dimension (100 in your case). - Use element-wise addition (
+=) to accumulate vectors correctly. - Add safeguards to avoid division by zero (in case none of the words are found in the model).
Modified Working Code
import re import csv import gensim import pandas as pd import numpy as np # Add this import def processing(kata): words = re.sub(r'([^\s\w]|_)', '', str(kata)) words = re.sub(r'[0-9]+', '', words) return words path = './model_terbaru/idwiki_word2vec_100.model' model = gensim.models.word2vec.Word2Vec.load(path) data = pd.read_csv('data/coba_data1.csv', encoding="ISO-8859-1") for index, row in data.iterrows(): # Initialize total as 100-dimensional zero array (matches your model's vector size) total = np.zeros(100) kalimat = row[0] valid_word_count = 0 # Track how many words were actually found in the model # Preprocess the sentence words = re.sub(r'([^\s\w]|_)', '', str(kalimat)) words = re.sub(r'[0-9]+', '', words) for word in words.split(): kata = word.lower() try: vector = model[kata] # No need for extra quotes here total += vector # Element-wise vector addition valid_word_count += 1 print(f"Vector for {kata}: {vector}") print(f"Running total: {total}") except KeyError: # Catch specific KeyError instead of bare except print(f"Word '{kata}' not found in model, skipping") pass print(f"Total valid words found: {valid_word_count}") # Calculate average only if we have valid words if valid_word_count > 0: rata = total / valid_word_count # Element-wise division print(f"Average vector: {rata}") # Optional: Write to CSV # with open('data/vector_training.csv', 'a', newline='') as ok: # writer = csv.writer(ok) # writer.writerow(rata.tolist()) # Convert numpy array to list for CSV else: print("No valid words found in this sentence, skipping average calculation")
Key Changes Explained
np.zeros(100): Initializestotalas a 100-dimensional zero array, matching the output of your Word2Vec model.total += vector: Properly accumulates each word's vector using element-wise addition (the correct way to sum vectors).valid_word_count: Tracks how many words were actually found in the model (instead of usinglen(words.split()), which counts all words including those not in the model). This gives a more accurate average.- Specific
KeyErrorcatch: Avoids catching unexpected errors with a bareexcept. - Division safeguard: Prevents division by zero when no words are found in the model.
内容的提问来源于stack exchange,提问作者Muhammad Rusli
相关产品推荐
相关产品推荐

