如何拆分DataFrame中含列表格式字符串的列并扩展对应行?
处理DataFrame中列表格式字符串列的拆分展开
现有一个DataFrame,其中actor列的值是列表格式的字符串,示例代码如下:
import pandas as pd actor = ["[Emil Eifrem,Hugo Weaving,Laurence Fishburne]"] title = ["The Matrix"] actors = pd.DataFrame(list(zip(actor, title)), columns = ["actor", "title"])
我们需要将actor列的列表字符串拆分为独立元素,同时让title列对应重复多次,最终得到如下结构的DataFrame:
actor = ["Emil Eifrem", "Hugo Weaving", "Laurence Fishburne"] title = ["The Matrix"] * 3 actors = pd.DataFrame(list(zip(actor, title)), columns = ["actor", "title"])
解决方案
通过两步操作即可实现需求:
- 将列表格式的字符串转换为真实的列表对象
- 使用
explode方法将列表列展开为多行,自动重复对应行的其他列值
完整代码
import pandas as pd actor = ["[Emil Eifrem,Hugo Weaving,Laurence Fishburne]"] title = ["The Matrix"] actors = pd.DataFrame(list(zip(actor, title)), columns = ["actor", "title"]) # 清洗字符串,转换为列表 actors['actor'] = actors['actor'].str.strip('[]').str.split(',') # 展开列表列,重置索引 actors = actors.explode('actor', ignore_index=True) # 查看最终结果 print(actors)
输出结果
actor title 0 Emil Eifrem The Matrix 1 Hugo Weaving The Matrix 2 Laurence Fishburne The Matrix
内容的提问来源于stack exchange,提问作者Salahuddin
相关产品推荐
相关产品推荐

