You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Cython中C方式创建的列表转Set慢于纯Python的原因及优化方案

优化Cython中char指针数组转Python Set的性能方案

Great question! Let's break down why there's such a big performance gap when converting these two types of lists to sets, and then dive into the fixes.

The Root Cause of the Performance Gap

When you create a list from a char* array without converting the C strings to Python strings upfront, the list ends up storing Cython-wrapped char* objects—not native Python str instances. When you pass this list to set(), Python has to perform a type conversion for every single element: it takes each char*, converts it to a Python str, and then adds it to the set. This per-element conversion adds massive overhead compared to the other approach, where you've already converted the strings to Python str before adding them to the list.

Fix 1: Convert C Strings to Python Strings Upfront (Most Straightforward)

The simplest and safest fix is to convert each char* to a Python str while building the list, not when converting to a set. Use the Cython-exposed Python API PyUnicode_FromString for efficient conversion:

from libc.string cimport strcpy
from cpython.string cimport PyUnicode_FromString

def make_list_from_char_ptrs(char** c_strings, int n):
    cdef list py_list = []
    cdef int i
    for i in range(n):
        # Convert C string to Python str BEFORE adding to the list
        py_list.append(PyUnicode_FromString(c_strings[i]))
    return py_list

Now your list contains native Python str objects, just like the list built by appending elements directly. When you pass this to set(), there's no extra conversion work needed—performance will match the faster approach perfectly.

Fix 2: Skip the List Entirely (Build the Set Directly)

If you don't need the intermediate list, you can build the Python set directly from the char* array. This eliminates the overhead of creating and iterating over a list first:

from cpython.set cimport PySet_New, PySet_Add
from cpython.string cimport PyUnicode_FromString

def make_set_from_char_ptrs(char** c_strings, int n):
    cdef object py_set = PySet_New(NULL)
    cdef int i
    cdef object py_str
    for i in range(n):
        py_str = PyUnicode_FromString(c_strings[i])
        PySet_Add(py_set, py_str)
    return py_set

This is even more efficient for large datasets, as you avoid copying all elements into a list before moving them to the set.

Optional: Avoid Memory Copies (For Advanced Users)

If you control the lifecycle of the C strings (e.g., they're allocated in a way that Python can safely take ownership), you can use PyUnicode_FromStringAndSize to create Python strings without copying the underlying bytes. Warning: This is risky—you must ensure the C string memory isn't freed while the Python string exists.

from libc.string cimport strlen
from cpython.string cimport PyUnicode_FromStringAndSize

def make_list_from_char_ptrs_no_copy(char** c_strings, int n):
    cdef list py_list = []
    cdef int i
    cdef Py_ssize_t str_len
    for i in range(n):
        str_len = strlen(c_strings[i])
        # Create Python string without copying the C string bytes
        py_str = PyUnicode_FromStringAndSize(c_strings[i], str_len)
        py_list.append(py_str)
    return py_list

Only use this if you're confident you can manage the C memory correctly—incorrect use will lead to crashes or memory corruption.

内容的提问来源于stack exchange,提问作者Ted Petrou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:22:43