You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Python循环:加速词向量匹配与CSV写入?

Optimize Python Script for Fast Vector Lookup and CSV Writing

Hey there! Let's fix that slow script—processing 100k rows can crawl if you're relying on inefficient lookups and row-by-row IO. Here's how to slash runtime from hours to minutes (or even seconds):

1. Replace DataFrame Lookups with a Dictionary (O(1) Access)

The biggest bottleneck here is almost certainly searching for words in your vectors DataFrame repeatedly. DataFrame lookups are O(n) per query, which adds up fast with 100k rows. Instead, convert your vectors into a hash map (dictionary) for instant lookups:

import pandas as pd
import numpy as np
import csv

# Load your data (adjust headers if your CSVs have them)
texts = pd.read_csv("texts.csv", sep='\t', header=None)
vectors_df = pd.read_csv("word_n_vectors.csv", sep='\t')

# Convert vectors to a dictionary: key = word, value = 100D numpy array
# Replace `vectors_df.columns[0]` with your actual word column name if needed
vector_dict = {}
for _, row in vectors_df.iterrows():
    word = row[vectors_df.columns[0]]
    vector = np.array(row[1:], dtype=np.float32)  # Use float32 to save memory/speed
    vector_dict[word] = vector

# Even faster one-liner alternative for numpy arrays:
vector_dict = {k: np.array(v, dtype=np.float32) for k, v in vectors_df.set_index(vectors_df.columns[0]).iterrows()}

2. Batch Process Rows (Avoid Row-by-Row Loops)

Instead of processing and writing one row at a time, handle all your text data in batches, then write everything at once. This cuts down on expensive IO operations.

Case 1: Each row in texts is a single word

If each line in texts.csv is one word that maps directly to a 100D vector:

# Process all rows in one go
output_rows = []
for word in texts.iloc[:, 0]:
    # Fallback to a zero vector if the word isn't in your vector dict (adjust as needed)
    vector = vector_dict.get(word, np.zeros(100, dtype=np.float32))
    output_rows.append(vector.tolist())

# Write all data to CSV in a single operation
with open('output.csv', 'w', newline='') as f:
    writer = csv.writer(f, delimiter='\t')
    writer.writerows(output_rows)

# Or use Pandas for even cleaner code (and fast IO):
output_df = pd.DataFrame(output_rows)
output_df.to_csv('output.csv', sep='\t', index=False, header=False)

Case 2: Each row in texts is a sequence of words (e.g., space-separated)

If you need to compute a combined vector (like average) for each row's word sequence:

def get_combined_vector(word_string):
    words = word_string.split()
    # Collect vectors for valid words, fallback to zeros for missing ones
    vectors = [vector_dict.get(word, np.zeros(100, dtype=np.float32)) for word in words]
    if not vectors:
        return np.zeros(100, dtype=np.float32)
    # Return average vector (replace with sum/concatenation if needed)
    return np.mean(vectors, axis=0)

# Apply to all rows (use swifter library to auto-accelerate this step!)
# Install swifter first: pip install swifter
import swifter
texts['combined_vector'] = texts.iloc[:, 0].swifter.apply(get_combined_vector)

# Expand the vector column into 100 separate columns and write
output_df = pd.DataFrame(texts['combined_vector'].tolist())
output_df.to_csv('output.csv', sep='\t', index=False, header=False)

3. Extra Speed Boosts

  • Specify dtypes when loading CSVs: Tell Pandas exactly what data types to expect to avoid slow type inference. For example:
    vector_dtypes = {vectors_df.columns[0]: str}
    vector_dtypes.update({col: np.float32 for col in vectors_df.columns[1:]})
    vectors_df = pd.read_csv("word_n_vectors.csv", sep='\t', dtype=vector_dtypes)
    
  • Chunk large datasets: If texts.csv is too big for memory, process it in chunks:
    chunk_size = 10000
    with open('output.csv', 'w', newline='') as f:
        writer = csv.writer(f, delimiter='\t')
        for chunk in pd.read_csv("texts.csv", sep='\t', header=None, chunksize=chunk_size):
            chunk_vectors = [vector_dict.get(word, np.zeros(100)) for word in chunk.iloc[:, 0]]
            writer.writerows([v.tolist() for v in chunk_vectors])
    
  • Avoid unnecessary copies: Use numpy arrays instead of lists for vectors to reduce memory overhead and speed up computations.

The core issue with your original script was likely the repeated slow lookups in the DataFrame. Switching to a dictionary eliminates that bottleneck, and batch IO cuts down on the overhead of writing to disk one row at a time.

内容的提问来源于stack exchange,提问作者Anon George

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:51:28