You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

阿拉伯文本词云显示异常求助:代码未改动但显示错误

阿拉伯文本词云显示异常问题

数月前使用代码生成阿拉伯文本词云一切正常,如今复用相同代码却出现文字错乱、排版异常的问题。尝试更换多种阿拉伯字体后,问题仍未解决。

所用代码如下:

import os
from collections import Counter
import codecs
from wordcloud import WordCloud  
import matplotlib.pyplot as plt          
# -- Arabic text dependencies
from arabic_reshaper import reshape      
from bidi.algorithm import get_display  
#arb_stopwords = set(nltk.corpus.stopwords.words("arabic"))

rtl = lambda w: get_display(reshape(f'{w}'))
d = os.path.dirname(__file__) if "__file__" in locals() else os.getcwd()
f = codecs.open(os.path.join(d,"/content/drive/MyDrive/file.txt"), 'r', 'utf-8')
text = f.read()
COUNTS = Counter(text.split())
counts = {rtl(k):v for k, v in COUNTS.most_common(300)}


font_file = "/content/drive/MyDrive/NotoNaskhArabic-Regular.ttf" 

wordcloud = WordCloud(font_path=font_file,
                      width = 1000, height = 1000,
                      background_color ='white',
                      #collocations=False,
                      #stopwords = arb_stopwords,
                      min_font_size = 10).generate_from_frequencies(counts)
plt.imshow(wordcloud, interpolation="bilinear")
plt.axis("off")
plt.show()
print(COUNTS)
解决方法
  • 检查依赖库版本:异常大概率是arabic_reshaper、python-bidi或wordcloud版本更新导致的兼容性问题。可以回退到之前正常运行的版本,比如执行:

    pip install arabic-reshaper==2.1.4 python-bidi==0.4.2 wordcloud==1.8.2.2
    

    也可以更新到最新稳定版,同时适配新API,比如新版arabic_reshaper需添加配置保留元音符号:

    from arabic_reshaper import Configuration, reshape
    reshaped_text = reshape(text, configuration=Configuration(delete_harakat=False))
    
  • 修正文本处理逻辑:

    • 阿拉伯语不能直接用split()按空格分词,会导致单词拆分错误,建议使用pyarabic或nltk的阿拉伯语分词工具替换;
    • 清洗文本中的特殊字符,避免编码干扰:
      import re
      text = re.sub(r'[^\u0600-\u06FF\s]', '', text)  # 仅保留阿拉伯语字符与空格
      
  • 调整词云参数:

    • 取消注释collocations=False,避免词云拼接错误词组;
    • 验证字体路径有效性,可尝试系统自带阿拉伯字体(如Linux的Amiri、Windows的Arial Unicode MS)测试。
  • 重构RTL转换逻辑:替换原lambda函数为更健壮的实现,避免字符串格式化问题:

    def rtl(text):
        reshaped_text = reshape(text)
        return get_display(reshaped_text)
    

    再用这个函数生成counts字典。

内容的提问来源于stack exchange,提问作者laah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 05:27:09