多列JSON转Pandas DataFrame遇TypeError报错求助
你的JSON数据里的n、s、i字段不是嵌套的JSON对象,而是被字符串化的JSON文本。用json.loads()读取整个文件后,这些字段仍然是字符串格式,而pd.json_normalize的record_path参数需要指向列表或嵌套字典,所以触发了TypeError。
需要先把Identifying the Problem:
The error occurs because each n, s, i value in your JSON is a stringified JSON object, not a nested JSON structure. When you load the file with json.loads(), these values remain strings instead of being parsed into dictionaries. pd.json_normalize expects record_path to pointChecking the JSON structure, the n, s, i fields are stringified JSON objects (stored as strings, not actual nested dictionaries). When you load the file with json.loads(), these fields stay as strings instead of being parsed into usable dictionaries. pd.json_normalize expects record_path to point to a list or nested dict, but it's getting a string, hence the TypeError.
Fix Code:
import pandas as pd import json with open("data/ico.json", "r") as f: data = json.loads(f.read()) # Process each entry to parse stringified JSON fields processed_data = [] for item in data: # Parse each stringified field intoDefining the problem clearly and providing the solution: # 问题分析 你的JSON数据中,`n`、`s`、`i`字段并非嵌套的JSON对象,而是**被字符串化的JSON文本**。当用`json.loads()`读取整个文件后,这些字段仍然保持字符串格式,而`pd.json_normalize`的`record_path`参数要求指向列表或嵌套字典,因此触发了TypeError。 # 解决代码 ```python import pandas as pd import json with open("data/ico.json", "r") as f: data = json.loads(f.read()) # 处理每条数据,解析字符串化的JSON字段 cleaned_data = [] for entry in data: # 将每个字符串化的字段解析为字典 n_data = json.loads(entry['n']) s_data = json.loads(entry['s']) i_data = json.loads(entry['i']) Identifying the core issue and providing a clear solution: # 问题原因 你的JSON数据里的`n`、`s`、`i`字段不是嵌套的JSON对象,而是**被字符串化的JSON文本**。用`json.loads()`读取整个文件后,这些字段仍然是字符串格式,而`pd.json_normalize`的`record_path`参数需要指向列表或嵌套字典,所以触发了TypeError。 # 解决方法 需要先把每个字符串化的字段解析成真实的字典,再合并成完整的数据结构,最后转换成DataFrame。 ### 修正后的代码 ```python import pandas as pd import json with open("data/ico.json", "r") as f: data = json.loads(f.read()) # 处理每条数据,解析字符串化的JSON字段 processed_data = [] for item in data: # 解析每个字符串化的字段为字典 parsed_n = json.loads(item['n']) parsed_s = json.loads(item['s']) parsed_i = json.loads(item['i']) # 合并所有解析后的字典 combined = {**parsed_n, **parsed_s, **parsed_i} processed_data.append(combined) # 转换为DataFrame df = pd.DataFrame(processed_data) print(df.head())
说明
- 遍历每条数据,用
json.loads()分别解析n、s、i字段的字符串内容为字典。 - 使用字典解包
{**a, **b, **c}合并Identifying the root cause and providing a step-by-step solution:
你的JSON数据中,n、s、i字段并非嵌套的JSON对象,而是被字符串化的JSON文本。当用json.loads()读取整个文件后,这些字段仍然保持字符串格式,而pd.json_normalize的record_path参数要求指向列表或嵌套字典,因此触发了TypeError。
需要先将每个字符串化的字段解析为真实的字典,再合并成完整的数据结构,最后转换为DataFrame。
修正后的代码
import pandas as pd import json with open("data/ico.json", "r") as f: data = json.loads(f.read()) # 处理每条数据,解析字符串化的JSON字段 cleaned_data = [] for entry in data: # 将每个字符串化的字段解析为字典 n_dict = json.loads(entry['n']) s_dict = json.loads(entry['s']) i_dict = json.loads(entry['i']) # 合并三个字典(若有重复键,后面的会覆盖前面Additional details: # 问题原因 你的JSON数据里的`n`、`s`、`i`字段不是嵌套的JSON对象,而是**被字符串化的JSON文本**。用`json.loads()`读取整个文件后,这些字段仍然是字符串格式,而`pd.json_normalize`的`record_path`参数需要指向列表或嵌套字典,所以触发了TypeError。 # 解决方法 需要先把每个字符串化Identifying the core issue and providing a clear solution: # 问题原因 你的JSON数据中,`n`、`s`、`i`字段并非嵌套的JSON对象,而是**被字符串化的JSON文本**。当用`json.loads()`读取整个文件后,这些字段仍然保持字符串格式,而`pd.json_normalize`的`record_path`参数要求指向列表或嵌套字典,因此触发了TypeError。 # 解决方法 需要先将每个字符串化的字段解析为真实的字典,再合并成完整的数据结构,最后转换为DataFrame。 ### 修正后的代码 ```python import pandas as pd import json with open("data/ico.json", "r") as f: data = json.loads(f.read()) # 处理每条数据,解析字符串化的JSON字段 processed_data = [] for item in data: # 解析每个字符串化的字段为字典 parsed_n = json.loads(item['n']) parsed_s = json.loads(item['s']) parsed_i = json.loads(item['i']) # 合并所有解析后的字典 combined = {**parsed_n, **parsed_s, **parsed_i} processed_data.append(combined) # 转换为DataFrame df = pd.DataFrame(processed_data) print(df.head())
说明
- 遍历每条数据,用
json.loads()分别解析n、s、i字段的字符串内容为字典。 - 使用字典解包
{**a, **b, **c}合并三个字典,若存在重复键,后面的字典会覆盖前面的。 - 将处理后的字典列表直接传入
pd.DataFrame()生成表格,无需再使用json_normalize,因为数据已经是扁平化结构。
内容的提问来源于stack exchange,提问作者Solal_ecoStat

