You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tabulate生成grid表格时多字节UTF-8字符单元格末尾出现空格

解决tabulate输出含Emoji表格时的多余空格问题

问题背景

使用Kaggle的热门Emoji数据集,通过csv模块读取数据后,用tabulate库的grid格式生成表格时,所有包含多字节UTF-8字符(如Emoji)的单元格末尾会出现多余空格,导致表格对齐混乱。

数据样本

Hex,Rank,Emoji,Year,Category,Subcategory,Name
\x{1F602},1,😂,2010,Smileys & Emotion,face-smiling,face with tears of joy
\x{2764 FE0F},2,❤️,2010,Smileys & Emotion,emotion,red heart
\x{1F923},3,🤣,2016,Smileys & Emotion,face-smiling,rolling on the floor laughing
\x{1F44D},4,👍,2010,People & Body,hand-fingers-closed,thumbs up
\x{1F62D},5,😭,2010,Smileys & Emotion,face-concerned,loudly crying face
\x{1F64F},6,🙏,2010,People & Body,hands,folded hands
\x{1F618},7,😘,2010,Smileys & Emotion,face-affection,face blowing a kiss
\x{1F970},8,🥰,2018,Smileys & Emotion,face-affection,smiling face with hearts
\x{1F60D},9,😍,2010,Smileys & Emotion,face-affection,smiling face with heart-eyes
\x{1F60A},10,😊,2010,Smileys & Emotion,face-smiling,smiling face with smiling eyes

原代码

import csv
from tabulate import tabulate as tb

top_n = 20

with open("emojis/emojis.csv", "r") as f:
    reader = [i for i in csv.reader(f)]
    print(tb([[j.replace("\\x", "") for j in i] for i in reader[:top_n + 1]], headers="firstrow", tablefmt="grid"))

问题原因

tabulate默认将每个Unicode字符视为1个显示宽度,但Emoji这类宽字符在终端中实际占2个字符宽度。tabulate计算单元格宽度时用了错误的宽度值,导致填充了多余的空格。

解决方案

使用wcwidth库来正确计算字符的终端显示宽度,修正tabulate的宽度计算逻辑。

步骤1:安装依赖

pip install wcwidth

步骤2:修正后的代码

import csv
from tabulate import tabulate as tb
from wcwidth import wcswidth

# 替换tabulate内部的文本宽度计算函数,改用wcswidth获取真实显示宽度
tb._text_width = wcswidth

top_n = 20

with open("emojis/emojis.csv", "r", encoding="utf-8") as f:
    reader = list(csv.reader(f))
    # 无需替换\x,csv读取时会自动处理UTF-8转义
    print(tb(reader[:top_n + 1], headers="firstrow", tablefmt="grid"))

备选方案(不修改库内部逻辑)

如果不想改动tabulate的内部函数,可以手动计算每列的真实宽度并调整单元格内容:

import csv
from tabulate import tabulate as tb
from wcwidth import wcswidth

def pad_text(text, target_width):
    actual_width = wcswidth(str(text))
    return str(text) + " " * (target_width - actual_width)

top_n = 20

with open("emojis/emojis.csv", "r", encoding="utf-8") as f:
    reader = list(csv.reader(f))
    headers = reader[0]
    data_rows = reader[1:top_n+1]

    # 计算每列的最大显示宽度
    col_max_widths = []
    for col_idx in range(len(headers)):
        all_col_values = [row[col_idx] for row in reader[:top_n+1]]
        max_width = max(wcswidth(str(val)) for val in all_col_values)
        col_max_widths.append(max_width)

    # 调整每行的单元格内容,确保宽度匹配
    adjusted_data = []
    for row in data_rows:
        adjusted_row = [pad_text(cell, col_max_widths[idx]) for idx, cell in enumerate(row)]
        adjusted_data.append(adjusted_row)

    print(tb(adjusted_data, headers=headers, tablefmt="grid", stralign="left"))

内容的提问来源于stack exchange,提问作者Baba20

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 01:30:55