如何清理DataFrame中genres列冗余内容,提取纯类型名称?
处理DataFrame中genres列冗余内容的解决方案
需求
移除DataFrame中genres列的冗余文字与符号,仅保留类型名称。
原始数据示例
| id | genres |
|---|---|
| 19995 | [{"id": 28, "name": "Action"}] |
| 285 | [{"id": 12, "name": "Adventure"}] |
尝试代码及错误
尝试代码:
json.loads(movies['genre'][0]).values()
触发错误:
AttributeError: 'list' object has no attribute 'values'
错误原因:json.loads()解析后得到的是列表(原始JSON是数组格式),列表没有values()方法,需要遍历列表内的字典提取name字段。
正确解决方法
场景1:genres列是JSON格式字符串
使用apply()结合列表推导式,解析JSON并提取所有name字段,用逗号拼接:
import json import pandas as pd movies['genres'] = movies['genres'].apply( lambda x: ','.join([item['name'] for item in json.loads(x)]) )
场景2:genres列已是Python列表(已完成JSON解析)
如果数据读取时已自动解析JSON,genres直接存储列表,可跳过json.loads():
movies['genres'] = movies['genres'].apply( lambda x: ','.join([item['name'] for item in x]) )
处理后期望结果
| id | genres |
|---|---|
| 19995 | Action,Adventure,Fantasy,Science Fiction |
| 285 | Adventure,Fantasy,Action |
内容的提问来源于stack exchange,提问作者Inoue
相关产品推荐
相关产品推荐

