You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JSON转Pandas DataFrame时文本显示异常问题排查

问题描述

我通过HTTP请求加载数据并保存为JSON文件,相关代码如下:

数据请求与处理代码

try:
    response = requests.get(url)
except requests.exceptions.ConnectionError as e:
    print("ConnectionError:", e)
try:
    data = json.loads(response.text)
except JSONDecodeError:
    if "Bad Gateway" in response.text:
        print("Bad Gateway, sleeping.")
    else:
        print("JSONDecodeError, data not loaded")
if data and data["response"]["status"] == 200:
    for m in data["messages"]:
        record = {}           
        record["id"] = m["id"]
        record["text"] = m["body"]
        tweets.append(record)

保存JSON文件代码

tweets是存储请求信息的字典列表,我用以下代码将其保存为JSON文件:

fileName = "stuff/{}/{}.json".format(symbol, Id)
with open(fileName, "w",encoding="utf-8") as f:
    json.dump(tweets, f,ensure_ascii=False)

加载为DataFrame代码

随后我将JSON文件加载为DataFrame:

with open('thefile.json', 'r', encoding='utf-8') as f:
    data = json.load(f)
df = pd.read_json(json.dumps(data), orient='records')

异常现象

JSON文件中的文本内容为:"i felt awful to not cash out the 145 puts on Thursday open",但在Jupyter Notebook中使用df.head()查看时,文本显示为:"𝑖𝑓𝑒𝑙𝑡𝑎𝑤𝑓𝑢𝑙𝑡𝑜𝑛𝑜𝑡𝑐𝑎𝑠ℎ𝑜𝑢𝑡𝑡ℎ𝑒145𝑝𝑢𝑡𝑠𝑜𝑛𝑇ℎ𝑢𝑟𝑠𝑑𝑎𝑦𝑜𝑝𝑒𝑛"。不过将DataFrame转为Numpy数组(a = df["text"].values)后打印文本,显示是正常的。

请问这是数据本身存在问题,还是df.head()的显示样式导致的异常?


问题分析与结论

这是Jupyter Notebook中pandas的显示渲染问题,数据本身没有问题。

原因如下:

  • 使用df.head()时,Jupyter会调用pandas的HTML渲染器格式化输出,部分场景下渲染规则会把普通英文字符识别渲染成数学斜体字符(即你看到的连体样式);
  • 转为Numpy数组后打印,是直接输出字符串原始内容,绕过了pandas的HTML渲染逻辑,因此显示正常。

如果想在Jupyter里正常查看DataFrame文本内容,可以尝试:

  • 用print(df.head())输出,以纯文本形式代替HTML渲染;
  • 调整pandas显示设置,禁用HTML渲染:
pd.set_option('display.notebook_repr_html', False)

内容的提问来源于stack exchange,提问作者aaronm012

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 01:13:25