安装django-pgroonga遇编码错误,求日文向量搜索解决方案
问题:Django/PostgreSQL环境安装django-pgroonga时出现编码错误,寻求日文字符向量搜索替代方案
错误详情
执行pip install django-pgroonga时触发编码错误,完整报错信息如下:
PS C:\JGRAM\JLPT> pip install django-pgroonga Collecting django-pgroonga Using cached django-pgroonga-0.0.1.tar.gz (3.7 kB) Preparing metadata (setup.py) ... error error: subprocess-exited-with-error × python setup.py egg_info did not run successfully. │ exit code: 1 ╰─> [10 lines of output] Traceback (most recent call last): File "<string>", line 2, in <module> File "<pip-setuptools-caller>", line 34, in <module> File "C:\Users\61458\AppData\Local\Temp\pip-install-21w4o7u8\django-pgroonga_87013717bf0e4bcca83db91a993082b4\setup.py", line 17, in <module> long_description=read('README.rst'), File "C:\Users\61458\AppData\Local\Temp\pip-install-21w4o7u8\django-pgroonga_87013717bf0e4bcca83db91a993082b4\setup.py", line 6, in read return open(os.path.join(os.path.dirname(__file__), fname)).read() File "C:\Users\61458\AppData\Local\Programs\Python\Python310\lib\encodings\cp1252.py", line 23, in decode return codecs.charmap_decode(input,self.errors,decoding_table)[0] **UnicodeDecodeError: 'charmap' codec can't decode byte 0x81 in position 112: character maps to <undefined>** [end of output] note: This error originates from a subprocess, and is likely not a problem with pip. error: metadata-generation-failed × Encountered error while generating package metadata. ╰─> See above for output. note: This is an issue with the package mentioned above, not pip. hint: See above for details. PS C:\JGRAM\JLPT>
用户已尝试更新cp1252.py、手动解压包放入site-packages等操作,均未解决问题,且此前pip安装其他模块正常。
问题原因
该错误确实是django-pgroonga包本身的问题:包内setup.py的read函数读取README.rst时未指定编码,Windows系统默认用cp1252编码读取含日文字符的文件,触发了解码失败。该包版本为0.0.1,属于早期测试版本,维护不完善。
临时解决办法(若仍想使用django-pgroonga)
- 手动下载包源码,修改
setup.py中的read函数,添加UTF-8编码参数:def read(fname): return open(os.path.join(os.path.dirname(__file__), fname), encoding='utf-8').read() - 进入源码目录执行
pip install .完成安装。
日文字符向量搜索替代方案
如果不想折腾该包,推荐以下工具:
- pgvector + Django集成:pgvector是PostgreSQL官方支持的向量扩展,搭配
django-pgvector包可快速与Django整合。针对日文字符,先通过MeCab、日文版Jieba等工具分词,再用日本语BERT、OpenAI Embedding等模型生成向量,存入pgvector字段即可实现向量搜索。 - Elasticsearch + Django:Elasticsearch自带日文分词器(kuromoji),对多语言支持完善,可同时实现日文字符的全文检索和向量搜索。通过
django-elasticsearch-dsl包与Django集成,适合复杂检索场景。 - 直接使用Pgroonga扩展:跳过
django-pgroonga包,直接在PostgreSQL中安装Pgroonga扩展,在Django中通过原生SQL或自定义模型方法调用Pgroonga的搜索功能,规避Python包的编码问题。
内容的提问来源于stack exchange,提问作者user19565014
相关产品推荐
相关产品推荐

