You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HTML爬取数据转NumPy遇序列元素错误,求内存优化及IO加速方案

Fixing NumPy Array Issues for Large-Scale Web Scraping Data

Let's break down your problem and fix it step by step:

Why You're Getting the "setting an array element with a sequence" Error

The core issue here is twofold:

  • Your SapList contains nested sequences (like [strLink, ProdName] or the sublists in ss), but NumPy arrays expect uniform, consistent data structures. When you try to convert this irregular list to an array, NumPy can't infer a valid dtype, triggering the error.
  • Using np.append in a loop is catastrophically inefficient—every call creates a brand new array by copying the entire existing dataset, which will tank performance and blow up memory usage as you scale to 50k entries.

Step 1: Restructure Your Data for NumPy

First, we need to standardize each shop's data into a format NumPy can handle cleanly. Since you have mixed data types (integers, strings, nested lists), a structured array is the perfect fit—it lets you define named fields with specific dtypes, like a table row.

Rework your eachshop function to return a single structured record instead of modifying global variables:

import numpy as np
from multiprocessing.pool import ThreadPool

def eachshop(args):
    # Unpack arguments from the enumerate() call
    row_num, sap_data = args
    # Replace these placeholders with your actual scraped values
    strLink = "https://example-shop.com/item1"
    ProdName = "Wireless Headphones"
    ProdCode = "WH-700"
    ProdH = "Premium"
    NewPrice = 199.99
    OldPrice = 249.99
    FileName = "scrape_batch_01"
    KompPrice = 189.99
    ss = [["id_001", "https://link1.com", "123 Main St"], ["id_002", "https://link2.com", "456 Oak Ave"]]

    # Convert nested lists to delimited strings (easier to store in structured arrays)
    link_name = f"{strLink}|{ProdName}"
    komp_entry = f"{FileName}#Komp!A1|{KompPrice}"
    sav_entry = f"{FileName}#Sav!A1|Sav"
    ss_str = "||".join([f"{item[0]}|{item[1]}|{item[2]}" for item in ss])

    # Return a tuple matching the structured dtype we'll define later
    return (
        row_num,
        sap_data,
        link_name,
        ProdCode,
        ProdH,
        NewPrice,
        OldPrice,
        komp_entry,
        sav_entry,
        ss_str
    )

Step 2: Efficiently Collect Data with ThreadPool

Instead of modifying a global NumPy array in each thread (which is also thread-unsafe!), collect all results from the thread pool first, then convert to a NumPy array in one go:

def makePool(cP, func, iters):
    all_records = []
    try:
        with ThreadPool(cP) as pool:
            # map() returns a list of all results from eachshop calls
            all_records = pool.map(func, enumerate(iters, start=2))
    except Exception as e:
        print(f'Pool Error: {str(e)}')
        raise
    return all_records

# Example usage: replace `your_scrape_iterable` with your actual list of URLs/items
scrape_results = makePool(4, eachshop, your_scrape_iterable)

# Define the structured dtype for our array (adjust string lengths to match your data)
dtype = [
    ('row_num', int),
    ('sap', 'U50'),
    ('link_name', 'U200'),
    ('prod_code', 'U50'),
    ('prod_h', 'U20'),
    ('new_price', float),
    ('old_price', float),
    ('komp_entry', 'U100'),
    ('sav_entry', 'U100'),
    ('ss_data', 'U1000')
]

# Convert the list of tuples to a structured NumPy array
list_all = np.array(scrape_results, dtype=dtype)

Step 3: Optimize Memory & IO Performance

Memory Savings

  • Using a structured array with specific dtypes (instead of an object array) cuts memory usage drastically. Fixed-length strings (U50) take far less space than storing Python list objects.
  • Ditching incremental np.append eliminates the overhead of repeated full-array copies.

Fast IO with NumPy

Once you have your structured array, use NumPy's built-in binary IO functions for lightning-fast saves/loads:

# Save to a compact binary file (fastest for large datasets)
np.save('scraped_data.npy', list_all)

# Load the array back later
loaded_data = np.load('scraped_data.npy')

# For human-readable output, use savetxt (binary is still better for performance)
np.savetxt(
    'scraped_data.csv', 
    list_all, 
    fmt='%s', 
    delimiter=',', 
    header=','.join([field[0] for field in dtype])
)

Key Notes for Your Use Case

  • If you need to keep nested lists intact (instead of converting to strings), use object dtype for those fields. Just note this won't save as much memory as structured dtypes:
    dtype = [
        # ... other fields ...
        ('ss_data', object)  # Stores Python lists directly
    ]
    
  • Thread safety: Using pool.map to collect results into a list avoids race conditions that would happen if multiple threads modified a global NumPy array simultaneously.
  • Pre-allocation: If you know the exact number of entries upfront, pre-allocate the structured array and fill it directly for even better efficiency:
    total_entries = len(your_scrape_iterable)
    list_all = np.empty(total_entries, dtype=dtype)
    # Fill entries with pool.imap_unordered for chunked processing
    

内容的提问来源于stack exchange,提问作者Dmitrij Holkin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:54:13