如何高效向Python 3函数传递大数据?
Hey Jared, great question—coming from C++, it makes total sense to want that fine-grained control over how data is passed around, especially when dealing with huge arrays where copies would tank performance and waste memory. Let’s break this down clearly:
首先:Python的参数传递是传对象引用,而非严格意义上的传值/传引用
Unlike C++ where you explicitly pick between pass-by-value and pass-by-reference, Python uses a "pass-by-object-reference" model. For your use case, this translates to:
- When you pass a mutable object (like a list, numpy array, or dictionary) to a function, you’re not sending a copy of the data—you’re sending a reference to the original object. No massive memory overhead, no time wasted copying gigabytes of data.
- Any modifications you make to the object’s contents (not reassigning the variable itself) will show up outside the function, just like with C++ references.
针对超大数组的两种实用场景
1. 使用Python内置list
Lists are mutable, so passing them to functions works exactly how you’d want to avoid copies:
- Tweaking elements of the list inside the function will change the original list.
- The only gotcha: if you reassign the list variable inside the function (e.g.,
arr = [new, values]), you’re creating a new object and breaking the link to the original—this won’t affect the external list.
Example code:
def update_large_list(arr): # 修改原数组的元素——外部会看到变化 arr[0] = 1000 arr[-1] = 9999 # 下面这行是重新赋值,不会影响外部的arr # arr = [1,2,3] huge_list = [i for i in range(1_000_000)] # 百万元素的大列表 update_large_list(huge_list) print(huge_list[0], huge_list[-1]) # 输出 1000 9999
2. 使用numpy数组(推荐处理数值超大数组)
If you’re working with numerical arrays (the most common case for "huge arrays"), numpy is your go-to tool. Numpy arrays are built for efficient memory usage and fast operations, and they behave almost exactly like C++ arrays with references:
- Passing a numpy array to a function never copies the data by default (unless you explicitly call
.copy()). - Any in-place modifications (like changing elements, scaling, or slicing with assignment) directly alter the original array.
Example code:
import numpy as np def process_large_numpy_arr(arr): # 原地修改数组——外部会立即看到变化 arr[::1000] = 0 # 每隔1000个元素设为0 arr *= 2 # 所有元素乘以2 huge_np_arr = np.arange(1_000_000) # 百万元素的numpy数组 process_large_numpy_arr(huge_np_arr) print(huge_np_arr[0], huge_np_arr[999], huge_np_arr[1000]) # 输出 0 1998 0
关键总结
- To avoid copying huge arrays: use mutable objects like
listor (preferably)numpy.ndarray—they’re passed by reference, so no duplicate data is created. - To keep modifications outside the function: modify the object’s contents (e.g.,
arr[index] = value) instead of reassigning the entire variable (e.g.,arr = new_array). - For numerical work, numpy is far more efficient than standard lists for large datasets, and its behavior aligns closely with what you’re used to in C++.
内容的提问来源于stack exchange,提问作者Jared

