Cython优化:检查列表单词是否为字符串子串
优化Cython子串检查性能的问题
初始实现与性能瓶颈
需求:遍历输入单词列表list_words,检查是否存在单词是输入字符串的子串,要求大小写不敏感(输入单词列表已预先转为小写)。
编写的初始Cython代码如下:
cpdef cy_check_any_word_is_substring(list_words, string): cdef unicode w cdef unicode s_lowered = string.lower() for w in list_words: if w in s_lowered: return True return False
使用示例:
# 所有list_words中的单词已转为小写 list_words = ['cat', 'dog', 'eat', 'seat'] input_string = 'The animal saw the Dog and started to make noises' # 应返回True cy_check_any_word_is_substring(list_words, input_string)
问题:代码注解后大部分内容呈黄色,说明存在大量Python交互,性能未达到预期,需要优化。
更新1:尝试C++容器方案报错
为提升性能,尝试改用C++的vector和string实现,编写代码如下:
from libcpp.vector cimport vector from libcpp.string cimport string cpdef cy_check_any_word_is_substring(vector[string] list_words,string string): s_lowered = string.lower() for w in list_words: if w in s_lowered: return True return False
编译时出现错误:
Invalid types for 'in' (string, Python object)
更新2:修改C++方案后性能反而下降
针对上述错误调整代码,改用C++ string的find方法实现子串检查:
from libcpp.vector cimport vector from libcpp.string cimport string,npos cdef bint cy_check_w_substring(string s_lowered, vector[string] list_words): cdef string w for w in list_words: if s_lowered.find(w) !=npos: return True return False cpdef cy3_check_any_word_is_substring(words_bytes, input_string): cdef bint result = False s_lowered = input_string.lower() result = cy_check_w_substring(bytes(s_lowered, 'utf8'), words_bytes) return result
使用时需要将原单词列表转为字节列表:
# 所有list_words中的单词已转为小写 list_words = ['cat', 'dog', 'eat', 'seat'] list_words_bytes = [bytes(w,'utf8') for w in list_words] input_string = 'The animal saw the Dog and started to make noises' # 应返回True cy3_check_any_word_is_substring(list_words_bytes, input_string)
但性能测试显示该方案表现更差:
%%timeit cy3_check_any_word_is_substring(list_words_bytes, input_string) #1.01 µs ± 3.16 ns per loop
对比初始方案的性能:
%%timeit cy_check_any_word_is_substring(list_words, input_string) #190 ns ± 0.773 ns per loop
原因分析:cy3_check_any_word_is_substring中把s_lowered转为bytes的操作就耗时145ns±0.55ns/循环,几乎等于初始方案的总耗时,直接拉低了整体性能。单独测试转字节操作的耗时:
%%timeit bytes(input_string, 'utf8') #145 ns ± 0.55 ns per loop
内容的提问来源于stack exchange,提问作者David Buchaca
相关产品推荐
相关产品推荐

