HTML爬取数据转NumPy遇序列元素错误,求内存优化及IO加速方案
Let's break down your problem and fix it step by step:
Why You're Getting the "setting an array element with a sequence" Error
The core issue here is twofold:
- Your
SapListcontains nested sequences (like[strLink, ProdName]or the sublists inss), but NumPy arrays expect uniform, consistent data structures. When you try to convert this irregular list to an array, NumPy can't infer a valid dtype, triggering the error. - Using
np.appendin a loop is catastrophically inefficient—every call creates a brand new array by copying the entire existing dataset, which will tank performance and blow up memory usage as you scale to 50k entries.
Step 1: Restructure Your Data for NumPy
First, we need to standardize each shop's data into a format NumPy can handle cleanly. Since you have mixed data types (integers, strings, nested lists), a structured array is the perfect fit—it lets you define named fields with specific dtypes, like a table row.
Rework your eachshop function to return a single structured record instead of modifying global variables:
import numpy as np from multiprocessing.pool import ThreadPool def eachshop(args): # Unpack arguments from the enumerate() call row_num, sap_data = args # Replace these placeholders with your actual scraped values strLink = "https://example-shop.com/item1" ProdName = "Wireless Headphones" ProdCode = "WH-700" ProdH = "Premium" NewPrice = 199.99 OldPrice = 249.99 FileName = "scrape_batch_01" KompPrice = 189.99 ss = [["id_001", "https://link1.com", "123 Main St"], ["id_002", "https://link2.com", "456 Oak Ave"]] # Convert nested lists to delimited strings (easier to store in structured arrays) link_name = f"{strLink}|{ProdName}" komp_entry = f"{FileName}#Komp!A1|{KompPrice}" sav_entry = f"{FileName}#Sav!A1|Sav" ss_str = "||".join([f"{item[0]}|{item[1]}|{item[2]}" for item in ss]) # Return a tuple matching the structured dtype we'll define later return ( row_num, sap_data, link_name, ProdCode, ProdH, NewPrice, OldPrice, komp_entry, sav_entry, ss_str )
Step 2: Efficiently Collect Data with ThreadPool
Instead of modifying a global NumPy array in each thread (which is also thread-unsafe!), collect all results from the thread pool first, then convert to a NumPy array in one go:
def makePool(cP, func, iters): all_records = [] try: with ThreadPool(cP) as pool: # map() returns a list of all results from eachshop calls all_records = pool.map(func, enumerate(iters, start=2)) except Exception as e: print(f'Pool Error: {str(e)}') raise return all_records # Example usage: replace `your_scrape_iterable` with your actual list of URLs/items scrape_results = makePool(4, eachshop, your_scrape_iterable) # Define the structured dtype for our array (adjust string lengths to match your data) dtype = [ ('row_num', int), ('sap', 'U50'), ('link_name', 'U200'), ('prod_code', 'U50'), ('prod_h', 'U20'), ('new_price', float), ('old_price', float), ('komp_entry', 'U100'), ('sav_entry', 'U100'), ('ss_data', 'U1000') ] # Convert the list of tuples to a structured NumPy array list_all = np.array(scrape_results, dtype=dtype)
Step 3: Optimize Memory & IO Performance
Memory Savings
- Using a structured array with specific dtypes (instead of an object array) cuts memory usage drastically. Fixed-length strings (
U50) take far less space than storing Python list objects. - Ditching incremental
np.appendeliminates the overhead of repeated full-array copies.
Fast IO with NumPy
Once you have your structured array, use NumPy's built-in binary IO functions for lightning-fast saves/loads:
# Save to a compact binary file (fastest for large datasets) np.save('scraped_data.npy', list_all) # Load the array back later loaded_data = np.load('scraped_data.npy') # For human-readable output, use savetxt (binary is still better for performance) np.savetxt( 'scraped_data.csv', list_all, fmt='%s', delimiter=',', header=','.join([field[0] for field in dtype]) )
Key Notes for Your Use Case
- If you need to keep nested lists intact (instead of converting to strings), use
objectdtype for those fields. Just note this won't save as much memory as structured dtypes:dtype = [ # ... other fields ... ('ss_data', object) # Stores Python lists directly ] - Thread safety: Using
pool.mapto collect results into a list avoids race conditions that would happen if multiple threads modified a global NumPy array simultaneously. - Pre-allocation: If you know the exact number of entries upfront, pre-allocate the structured array and fill it directly for even better efficiency:
total_entries = len(your_scrape_iterable) list_all = np.empty(total_entries, dtype=dtype) # Fill entries with pool.imap_unordered for chunked processing
内容的提问来源于stack exchange,提问作者Dmitrij Holkin

