如何用Python将Beautiful Soup爬取的不可读JSON格式化并提取数据?
Python处理爬取的JSON数据:格式化与指定数据提取
一、将JSON数据转换为可读格式并保存
你爬取到的原始JSON是压缩后的单行格式,用Python的json模块可以轻松转成带缩进的可读格式,同时处理可能的JSON语法问题:
import json # 替换为你从BeautifulSoup获取到的原始JSON字符串 raw_json_str = '{"props":{"pageProps":{"profile":{"user_id":"588e3c6fdd927a7cfad7beed","username":"5OC","outfit":[{"item_id":"body-flesh","name":"","rarity":"","active_palette":3,"parts":[],"colors":{"dependent_colors":[],"palettes":[]}}]}}}' # 解析JSON字符串为Python字典 try: data_dict = json.loads(raw_json_str) except json.JSONDecodeError as e: print(f"JSON解析失败: {e}") # 若原始数据存在语法缺失(比如你提供的示例少了闭合括号),需手动补全后再解析 # 格式化并写入文件 with open("formatted_data.json", "w", encoding="utf-8") as file: json.dump(data_dict, file, indent=2, ensure_ascii=False)
indent=2:设置缩进为2个空格,让JSON结构分层显示ensure_ascii=False:保留非ASCII字符(如中文)的原始格式- 用
try-except捕获解析错误,避免因原始JSON语法不完整导致程序崩溃
二、提取指定数据并输出
从解析后的字典中按层级提取目标字段(比如user_id和username),推荐用get()方法避免键不存在时报错:
# 逐层提取profile数据 profile = data_dict.get("props", {}).get("pageProps", {}).get("profile", {}) # 提取指定字段 user_id = profile.get("user_id") username = profile.get("username") # 打印输出 print(f"用户ID: {user_id}") print(f"用户名: {username}") # 也可将提取结果保存为单独的JSON文件 extracted_data = { "user_id": user_id, "username": username } with open("extracted_profile.json", "w", encoding="utf-8") as file: json.dump(extracted_data, file, indent=2, ensure_ascii=False)
内容的提问来源于stack exchange,提问作者yahyaahmed8989
相关产品推荐
相关产品推荐

