使用tabulate生成grid表格时多字节UTF-8字符单元格末尾出现空格
解决tabulate输出含Emoji表格时的多余空格问题
问题背景
使用Kaggle的热门Emoji数据集,通过csv模块读取数据后,用tabulate库的grid格式生成表格时,所有包含多字节UTF-8字符(如Emoji)的单元格末尾会出现多余空格,导致表格对齐混乱。
数据样本
Hex,Rank,Emoji,Year,Category,Subcategory,Name \x{1F602},1,😂,2010,Smileys & Emotion,face-smiling,face with tears of joy \x{2764 FE0F},2,❤️,2010,Smileys & Emotion,emotion,red heart \x{1F923},3,🤣,2016,Smileys & Emotion,face-smiling,rolling on the floor laughing \x{1F44D},4,👍,2010,People & Body,hand-fingers-closed,thumbs up \x{1F62D},5,😭,2010,Smileys & Emotion,face-concerned,loudly crying face \x{1F64F},6,🙏,2010,People & Body,hands,folded hands \x{1F618},7,😘,2010,Smileys & Emotion,face-affection,face blowing a kiss \x{1F970},8,🥰,2018,Smileys & Emotion,face-affection,smiling face with hearts \x{1F60D},9,😍,2010,Smileys & Emotion,face-affection,smiling face with heart-eyes \x{1F60A},10,😊,2010,Smileys & Emotion,face-smiling,smiling face with smiling eyes
原代码
import csv from tabulate import tabulate as tb top_n = 20 with open("emojis/emojis.csv", "r") as f: reader = [i for i in csv.reader(f)] print(tb([[j.replace("\\x", "") for j in i] for i in reader[:top_n + 1]], headers="firstrow", tablefmt="grid"))
问题原因
tabulate默认将每个Unicode字符视为1个显示宽度,但Emoji这类宽字符在终端中实际占2个字符宽度。tabulate计算单元格宽度时用了错误的宽度值,导致填充了多余的空格。
解决方案
使用wcwidth库来正确计算字符的终端显示宽度,修正tabulate的宽度计算逻辑。
步骤1:安装依赖
pip install wcwidth
步骤2:修正后的代码
import csv from tabulate import tabulate as tb from wcwidth import wcswidth # 替换tabulate内部的文本宽度计算函数,改用wcswidth获取真实显示宽度 tb._text_width = wcswidth top_n = 20 with open("emojis/emojis.csv", "r", encoding="utf-8") as f: reader = list(csv.reader(f)) # 无需替换\x,csv读取时会自动处理UTF-8转义 print(tb(reader[:top_n + 1], headers="firstrow", tablefmt="grid"))
备选方案(不修改库内部逻辑)
如果不想改动tabulate的内部函数,可以手动计算每列的真实宽度并调整单元格内容:
import csv from tabulate import tabulate as tb from wcwidth import wcswidth def pad_text(text, target_width): actual_width = wcswidth(str(text)) return str(text) + " " * (target_width - actual_width) top_n = 20 with open("emojis/emojis.csv", "r", encoding="utf-8") as f: reader = list(csv.reader(f)) headers = reader[0] data_rows = reader[1:top_n+1] # 计算每列的最大显示宽度 col_max_widths = [] for col_idx in range(len(headers)): all_col_values = [row[col_idx] for row in reader[:top_n+1]] max_width = max(wcswidth(str(val)) for val in all_col_values) col_max_widths.append(max_width) # 调整每行的单元格内容,确保宽度匹配 adjusted_data = [] for row in data_rows: adjusted_row = [pad_text(cell, col_max_widths[idx]) for idx, cell in enumerate(row)] adjusted_data.append(adjusted_row) print(tb(adjusted_data, headers=headers, tablefmt="grid", stralign="left"))
内容的提问来源于stack exchange,提问作者Baba20
相关产品推荐
相关产品推荐

