Pandas DataFrame应用样式后round方法失效的问题排查
Pandas Style格式化小数失效问题解决
问题场景
给DataFrame添加行高亮和表格样式后,尝试对start_hyp_minus_ref、dur_hyp_minus_ref列保留两位小数,使用.style.format()或DataFrame.round()方法均无效;但移除样式代码后,小数格式化功能正常。
原代码
import pandas as pd def printSegments(segments_ref, segments_hyp, timestamps_ref, timestamps_hyp): def highlight_unequal_rows(row): if len(row['reference'].split()) != len(row['hypothesis'].split()): return ['background-color: red'] * len(row) return [''] * len(row) start_diff = [th["start"] - tr["start"] for th, tr in zip(timestamps_hyp, timestamps_ref)] dur_diff = [th["dur"] - tr["dur"] for th, tr in zip(timestamps_hyp, timestamps_ref)] df = pd.DataFrame(dict(reference=segments_ref, hypothesis=segments_hyp, start_hyp_minus_ref=start_diff, dur_hyp_minus_ref=dur_diff)) pd.set_option('display.max_colwidth', None) df = df.style.apply(highlight_unequal_rows, axis=1) df = df.set_table_styles( [ {'selector': 'td', 'props': [('direction', 'ltr'), ('text-align', 'left')]}, {'selector': 'th', 'props': [('direction', 'ltr'), ('text-align', 'left')]} ], overwrite=False ) return df
尝试过的无效方法
- 错误调用Styler格式化方法(未对已有Styler对象操作):
f = {'start_hyp_minus_ref':'{:.2f}'} df.style.format(f).bar(subset='start_hyp_minus_ref')
- 转换为Styler后尝试用DataFrame方法修改列值:
df['start_hyp_minus_ref'] = df['start_hyp_minus_ref'].round(2) df['dur_hyp_minus_ref'] = df['dur_hyp_minus_ref'].round(2)
解决方案
问题核心是操作顺序错误:一旦将DataFrame转换为Styler对象,就无法再用DataFrame的方法修改数据;且格式化操作需要在样式、高亮之前执行,或直接在Styler对象上链式调用格式化方法。
修改后的代码:
import pandas as pd def printSegments(segments_ref, segments_hyp, timestamps_ref, timestamps_hyp): def highlight_unequal_rows(row): if len(row['reference'].split()) != len(row['hypothesis'].split()): return ['background-color: red'] * len(row) return [''] * len(row) start_diff = [th["start"] - tr["start"] for th, tr in zip(timestamps_hyp, timestamps_ref)] dur_diff = [th["dur"] - tr["dur"] for th, tr in zip(timestamps_hyp, timestamps_ref)] df = pd.DataFrame(dict(reference=segments_ref, hypothesis=segments_hyp, start_hyp_minus_ref=start_diff, dur_hyp_minus_ref=dur_diff)) pd.set_option('display.max_colwidth', None) # 先执行小数格式化,再应用高亮和表格样式 styler = df.style.format({ 'start_hyp_minus_ref': '{:.2f}', 'dur_hyp_minus_ref': '{:.2f}' }) styler = styler.apply(highlight_unequal_rows, axis=1) styler = styler.set_table_styles( [ {'selector': 'td', 'props': [('direction', 'ltr'), ('text-align', 'left')]}, {'selector': 'th', 'props': [('direction', 'ltr'), ('text-align', 'left')]} ], overwrite=False ) return styler
也可以用链式调用简化写法:
# 链式调用版本 styler = df.style.format({ 'start_hyp_minus_ref': '{:.2f}', 'dur_hyp_minus_ref': '{:.2f}' }).apply(highlight_unequal_rows, axis=1).set_table_styles( [ {'selector': 'td', 'props': [('direction', 'ltr'), ('text-align', 'left')]}, {'selector': 'th', 'props': [('direction', 'ltr'), ('text-align', 'left')]} ], overwrite=False )
测试用例
printSegments(["I went to the school"], ["I want to the school"], [{"start": 0.69, "dur": 1.5}], [{"start": 0.79, "dur": 1.3}])
运行后,start_hyp_minus_ref会显示0.10,dur_hyp_minus_ref显示-0.20,同时因两行单词数量一致,不会触发红色高亮。
内容的提问来源于stack exchange,提问作者Abdallah Barghouti
相关产品推荐
相关产品推荐

