Python如何导出含嵌套Comment数组的Post集合至CSV/Excel/JSON
如何将嵌套的Post/Comment对象结构导出为CSV、Excel或JSON?
首先是你定义的Python数据结构:
Contents = []; class Post: def __init__(self, poster,poster_url,post_text, post_images, post_comments): self.poster = poster self.poster_url = poster_url self.post_text = post_text self.post_images = post_images self.post_comments = post_comments class Comment: def __init__(self, poster,poster_url,post_text, post_images): self.poster = poster self.poster_url = poster_url self.post_text = post_text self.post_images = post_images
以下是针对三种格式的具体实现方案:
一、导出为JSON格式
JSON原生支持嵌套结构,只需将对象转换为字典后,用json模块写入文件即可。给Post和Comment类添加to_dict()方法,实现对象到字典的转换:
import json class Post: def __init__(self, poster,poster_url,post_text, post_images, post_comments): self.poster = poster self.poster_url = poster_url self.post_text = post_text self.post_images = post_images self.post_comments = post_comments def to_dict(self): return { "poster": self.poster, "poster_url": self.poster_url, "post_text": self.post_text, "post_images": self.post_images, "post_comments": [comment.to_dict() for comment in self.post_comments] } class Comment: def __init__(self, poster,poster_url,post_text, post_images): self.poster = poster self.poster_url = poster_url self.post_text = post_text self.post_images = post_images def to_dict(self): return { "poster": self.poster, "poster_url": self.poster_url, "post_text": self.post_text, "post_images": self.post_images } # 执行导出 with open("posts.json", "w", encoding="utf-8") as f: json.dump([post.to_dict() for post in Contents], f, ensure_ascii=False, indent=2)
二、导出为CSV格式
CSV是扁平结构,需处理嵌套的评论数据,有两种常用方式:
方式1:每条评论对应一行(推荐)
每行记录携带所属帖子的完整信息,便于后续数据分析:
import csv # 定义CSV表头,包含帖子+评论的所有字段 headers = [ "post_poster", "post_poster_url", "post_text", "post_images", "comment_poster", "comment_poster_url", "comment_text", "comment_images" ] with open("posts_comments.csv", "w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=headers) writer.writeheader() for post in Contents: # 处理无评论的帖子 if not post.post_comments: writer.writerow({ "post_poster": post.poster, "post_poster_url": post.poster_url, "post_text": post.post_text, "post_images": ",".join(post.post_images) if isinstance(post.post_images, list) else post.post_images, "comment_poster": "", "comment_poster_url": "", "comment_text": "", "comment_images": "" }) continue # 每条评论对应一行,重复帖子信息 for comment in post.post_comments: writer.writerow({ "post_poster": post.poster, "post_poster_url": post.poster_url, "post_text": post.post_text, "post_images": ",".join(post.post_images) if isinstance(post.post_images, list) else post.post_images, "comment_poster": comment.poster, "comment_poster_url": comment.poster_url, "comment_text": comment.post_text, "comment_images": ",".join(comment.post_images) if isinstance(comment.post_images, list) else comment.post_images })
方式2:将评论序列化为字符串存入单元格
把所有评论转成JSON字符串,放在帖子行的单独单元格中:
import csv import json headers = ["poster", "poster_url", "post_text", "post_images", "post_comments"] with open("posts_flat.csv", "w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=headers) writer.writeheader() for post in Contents: writer.writerow({ "poster": post.poster, "poster_url": post.poster_url, "post_text": post.post_text, "post_images": ",".join(post.post_images) if isinstance(post.post_images, list) else post.post_images, "post_comments": json.dumps([comment.to_dict() for comment in post.post_comments], ensure_ascii=False) })
三、导出为Excel格式
使用openpyxl库处理Excel,同样推荐扁平化存储评论数据,先安装依赖:
pip install openpyxl
实现代码:
from openpyxl import Workbook wb = Workbook() ws = wb.active # 设置表头 headers = [ "帖子发布者", "帖子发布者链接", "帖子内容", "帖子图片", "评论发布者", "评论发布者链接", "评论内容", "评论图片" ] ws.append(headers) for post in Contents: if not post.post_comments: # 写入无评论的帖子 ws.append([ post.poster, post.poster_url, post.post_text, ",".join(post.post_images) if isinstance(post.post_images, list) else post.post_images, "", "", "", "" ]) continue # 每条评论对应一行记录 for comment in post.post_comments: ws.append([ post.poster, post.poster_url, post.post_text, ",".join(post.post_images) if isinstance(post.post_images, list) else post.post_images, comment.poster, comment.poster_url, comment.post_text, ",".join(comment.post_images) if isinstance(comment.post_images, list) else comment.post_images ]) # 保存文件 wb.save("posts_comments.xlsx")
内容的提问来源于stack exchange,提问作者Mst. Murshida Khnom
相关产品推荐
相关产品推荐

