从PushShift API的JSON转CSV时遭遇TypeError问题求助
解决PushShift API JSON转CSV的TypeError问题
你这是踩了两个常见的小坑,咱们一步步捋清楚问题出在哪:
错误原因分析
你遇到的TypeError: byte indices must be integers or slices, not str,核心问题是循环遍历的对象完全不对:
- 要么是你误把
page(requests返回的Response对象)当成了JSON数据来遍历——遍历Response对象会逐个读取字节,自然没法用字符串索引。 - 要么是你直接遍历了
page_json这个外层字典——PushShift的返回结构是{"data": [评论数组], ...},page_json是个字典,遍历它只能拿到"data"、"metadata"这些键字符串,同样没法用["data"]索引。
另外你的代码还有两个可以优化的细节:
- 不用手动
json.loads(page.text),requests的Response对象自带.json()方法,更简洁安全。 - 直接用
open()不做上下文管理容易漏关文件,建议用with语句自动处理。
修正后的完整代码
import requests import csv # 请求PushShift API url = 'https://api.pushshift.io/reddit/comment/search/?subreddit=science&filter=parent_id,id,author,created_utc,subreddit,body,score,permalink' response = requests.get(url) # 直接解析JSON响应 api_data = response.json() # 用with语句管理CSV文件,自动处理打开/关闭 with open("test.csv", 'w+', newline='', encoding='utf-8') as csv_file: writer = csv.writer(csv_file) # 写入CSV表头 writer.writerow(["id", "parent_id", "author", "created_utc", "subreddit", "body", "score"]) # 遍历API返回的评论数据数组(注意是api_data["data"],不是api_data本身) for comment in api_data["data"]: # 用get方法避免字段缺失导致的KeyError(比如有些评论可能没有score字段) row_content = [ comment.get("id", ""), comment.get("parent_id", ""), comment.get("author", ""), comment.get("created_utc", ""), comment.get("subreddit", ""), comment.get("body", ""), comment.get("score", "") ] writer.writerow(row_content)
关键修正点
- 遍历正确的对象:必须遍历
api_data["data"],这才是包含所有评论的数组。 - 安全的字段获取:用
comment.get(key, 默认值)替代直接索引,避免某些评论缺失字段时抛出KeyError。 - 编码与文件管理:添加
encoding='utf-8'确保中文/特殊字符正常写入,with语句自动关闭文件。
内容的提问来源于stack exchange,提问作者dhrice
相关产品推荐
相关产品推荐

