如何优化Python循环:加速词向量匹配与CSV写入?
Hey there! Let's fix that slow script—processing 100k rows can crawl if you're relying on inefficient lookups and row-by-row IO. Here's how to slash runtime from hours to minutes (or even seconds):
1. Replace DataFrame Lookups with a Dictionary (O(1) Access)
The biggest bottleneck here is almost certainly searching for words in your vectors DataFrame repeatedly. DataFrame lookups are O(n) per query, which adds up fast with 100k rows. Instead, convert your vectors into a hash map (dictionary) for instant lookups:
import pandas as pd import numpy as np import csv # Load your data (adjust headers if your CSVs have them) texts = pd.read_csv("texts.csv", sep='\t', header=None) vectors_df = pd.read_csv("word_n_vectors.csv", sep='\t') # Convert vectors to a dictionary: key = word, value = 100D numpy array # Replace `vectors_df.columns[0]` with your actual word column name if needed vector_dict = {} for _, row in vectors_df.iterrows(): word = row[vectors_df.columns[0]] vector = np.array(row[1:], dtype=np.float32) # Use float32 to save memory/speed vector_dict[word] = vector # Even faster one-liner alternative for numpy arrays: vector_dict = {k: np.array(v, dtype=np.float32) for k, v in vectors_df.set_index(vectors_df.columns[0]).iterrows()}
2. Batch Process Rows (Avoid Row-by-Row Loops)
Instead of processing and writing one row at a time, handle all your text data in batches, then write everything at once. This cuts down on expensive IO operations.
Case 1: Each row in texts is a single word
If each line in texts.csv is one word that maps directly to a 100D vector:
# Process all rows in one go output_rows = [] for word in texts.iloc[:, 0]: # Fallback to a zero vector if the word isn't in your vector dict (adjust as needed) vector = vector_dict.get(word, np.zeros(100, dtype=np.float32)) output_rows.append(vector.tolist()) # Write all data to CSV in a single operation with open('output.csv', 'w', newline='') as f: writer = csv.writer(f, delimiter='\t') writer.writerows(output_rows) # Or use Pandas for even cleaner code (and fast IO): output_df = pd.DataFrame(output_rows) output_df.to_csv('output.csv', sep='\t', index=False, header=False)
Case 2: Each row in texts is a sequence of words (e.g., space-separated)
If you need to compute a combined vector (like average) for each row's word sequence:
def get_combined_vector(word_string): words = word_string.split() # Collect vectors for valid words, fallback to zeros for missing ones vectors = [vector_dict.get(word, np.zeros(100, dtype=np.float32)) for word in words] if not vectors: return np.zeros(100, dtype=np.float32) # Return average vector (replace with sum/concatenation if needed) return np.mean(vectors, axis=0) # Apply to all rows (use swifter library to auto-accelerate this step!) # Install swifter first: pip install swifter import swifter texts['combined_vector'] = texts.iloc[:, 0].swifter.apply(get_combined_vector) # Expand the vector column into 100 separate columns and write output_df = pd.DataFrame(texts['combined_vector'].tolist()) output_df.to_csv('output.csv', sep='\t', index=False, header=False)
3. Extra Speed Boosts
- Specify dtypes when loading CSVs: Tell Pandas exactly what data types to expect to avoid slow type inference. For example:
vector_dtypes = {vectors_df.columns[0]: str} vector_dtypes.update({col: np.float32 for col in vectors_df.columns[1:]}) vectors_df = pd.read_csv("word_n_vectors.csv", sep='\t', dtype=vector_dtypes) - Chunk large datasets: If
texts.csvis too big for memory, process it in chunks:chunk_size = 10000 with open('output.csv', 'w', newline='') as f: writer = csv.writer(f, delimiter='\t') for chunk in pd.read_csv("texts.csv", sep='\t', header=None, chunksize=chunk_size): chunk_vectors = [vector_dict.get(word, np.zeros(100)) for word in chunk.iloc[:, 0]] writer.writerows([v.tolist() for v in chunk_vectors]) - Avoid unnecessary copies: Use numpy arrays instead of lists for vectors to reduce memory overhead and speed up computations.
The core issue with your original script was likely the repeated slow lookups in the DataFrame. Switching to a dictionary eliminates that bottleneck, and batch IO cuts down on the overhead of writing to disk one row at a time.
内容的提问来源于stack exchange,提问作者Anon George

