You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何pandas read_html自动移除小数分隔符?表格数值格式异常求助

解决Pandas read_html数值解析错误(逗号小数点变字符串)的问题

直接修改你调用read_html的代码,加上小数和千位分隔符的参数,就能让Pandas正确识别数值类型:

table = soup.find("table", attrs={"id": "stats_shooting"}) 
df = pd.read_html(str(table), decimal=',', thousands='.')[0]

关键参数说明

  • decimal=',':目标网站用逗号作为小数点分隔符(比如0,62),这个参数告诉Pandas按此规则解析小数,避免把0,62识别成字符串062。
  • thousands='.':网站里的大数值用句号做千位分隔(比如1.234代表1234),这个参数能让Pandas正确解析这类数值。

若仍有列是字符串的处理方法

如果个别列还是没自动转为数值类型,可手动强制转换,以xG列为例:

df['xG'] = pd.to_numeric(df['xG'], errors='coerce', decimal=',')
  • errors='coerce'会把无法转换的内容转为NaN,方便后续数据清理;decimal=','再次明确小数点格式。

新手简化方案

其实不需要先用Selenium+BeautifulSoup定位表格,直接用read_html指定目标URL和表格ID更简洁:

df = pd.read_html('https://fbref.com/it/comp/11/shooting/Statistiche-di-Serie-A', 
                  attrs={"id": "stats_shooting"}, 
                  decimal=',', 
                  thousands='.')[0]

这样省去了额外的网页解析步骤,对新手更友好,解析精度也更高。

内容的提问来源于stack exchange,提问作者Giorgio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 03:12:47