Cython中C方式创建的列表转Set慢于纯Python的原因及优化方案
Great question! Let's break down why there's such a big performance gap when converting these two types of lists to sets, and then dive into the fixes.
The Root Cause of the Performance Gap
When you create a list from a char* array without converting the C strings to Python strings upfront, the list ends up storing Cython-wrapped char* objects—not native Python str instances. When you pass this list to set(), Python has to perform a type conversion for every single element: it takes each char*, converts it to a Python str, and then adds it to the set. This per-element conversion adds massive overhead compared to the other approach, where you've already converted the strings to Python str before adding them to the list.
Fix 1: Convert C Strings to Python Strings Upfront (Most Straightforward)
The simplest and safest fix is to convert each char* to a Python str while building the list, not when converting to a set. Use the Cython-exposed Python API PyUnicode_FromString for efficient conversion:
from libc.string cimport strcpy from cpython.string cimport PyUnicode_FromString def make_list_from_char_ptrs(char** c_strings, int n): cdef list py_list = [] cdef int i for i in range(n): # Convert C string to Python str BEFORE adding to the list py_list.append(PyUnicode_FromString(c_strings[i])) return py_list
Now your list contains native Python str objects, just like the list built by appending elements directly. When you pass this to set(), there's no extra conversion work needed—performance will match the faster approach perfectly.
Fix 2: Skip the List Entirely (Build the Set Directly)
If you don't need the intermediate list, you can build the Python set directly from the char* array. This eliminates the overhead of creating and iterating over a list first:
from cpython.set cimport PySet_New, PySet_Add from cpython.string cimport PyUnicode_FromString def make_set_from_char_ptrs(char** c_strings, int n): cdef object py_set = PySet_New(NULL) cdef int i cdef object py_str for i in range(n): py_str = PyUnicode_FromString(c_strings[i]) PySet_Add(py_set, py_str) return py_set
This is even more efficient for large datasets, as you avoid copying all elements into a list before moving them to the set.
Optional: Avoid Memory Copies (For Advanced Users)
If you control the lifecycle of the C strings (e.g., they're allocated in a way that Python can safely take ownership), you can use PyUnicode_FromStringAndSize to create Python strings without copying the underlying bytes. Warning: This is risky—you must ensure the C string memory isn't freed while the Python string exists.
from libc.string cimport strlen from cpython.string cimport PyUnicode_FromStringAndSize def make_list_from_char_ptrs_no_copy(char** c_strings, int n): cdef list py_list = [] cdef int i cdef Py_ssize_t str_len for i in range(n): str_len = strlen(c_strings[i]) # Create Python string without copying the C string bytes py_str = PyUnicode_FromStringAndSize(c_strings[i], str_len) py_list.append(py_str) return py_list
Only use this if you're confident you can manage the C memory correctly—incorrect use will lead to crashes or memory corruption.
内容的提问来源于stack exchange,提问作者Ted Petrou

