You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从PushShift API的JSON转CSV时遭遇TypeError问题求助

解决PushShift API JSON转CSV的TypeError问题

你这是踩了两个常见的小坑,咱们一步步捋清楚问题出在哪:

错误原因分析

你遇到的TypeError: byte indices must be integers or slices, not str,核心问题是循环遍历的对象完全不对:

  • 要么是你误把page(requests返回的Response对象)当成了JSON数据来遍历——遍历Response对象会逐个读取字节,自然没法用字符串索引。
  • 要么是你直接遍历了page_json这个外层字典——PushShift的返回结构是{"data": [评论数组], ...},page_json是个字典,遍历它只能拿到"data"、"metadata"这些键字符串,同样没法用["data"]索引。

另外你的代码还有两个可以优化的细节:

  1. 不用手动json.loads(page.text),requests的Response对象自带.json()方法,更简洁安全。
  2. 直接用open()不做上下文管理容易漏关文件,建议用with语句自动处理。

修正后的完整代码

import requests
import csv

# 请求PushShift API
url = 'https://api.pushshift.io/reddit/comment/search/?subreddit=science&filter=parent_id,id,author,created_utc,subreddit,body,score,permalink'
response = requests.get(url)
# 直接解析JSON响应
api_data = response.json()

# 用with语句管理CSV文件,自动处理打开/关闭
with open("test.csv", 'w+', newline='', encoding='utf-8') as csv_file:
    writer = csv.writer(csv_file)
    # 写入CSV表头
    writer.writerow(["id", "parent_id", "author", "created_utc", "subreddit", "body", "score"])
    
    # 遍历API返回的评论数据数组(注意是api_data["data"],不是api_data本身)
    for comment in api_data["data"]:
        # 用get方法避免字段缺失导致的KeyError(比如有些评论可能没有score字段)
        row_content = [
            comment.get("id", ""),
            comment.get("parent_id", ""),
            comment.get("author", ""),
            comment.get("created_utc", ""),
            comment.get("subreddit", ""),
            comment.get("body", ""),
            comment.get("score", "")
        ]
        writer.writerow(row_content)

关键修正点

  1. 遍历正确的对象:必须遍历api_data["data"],这才是包含所有评论的数组。
  2. 安全的字段获取:用comment.get(key, 默认值)替代直接索引,避免某些评论缺失字段时抛出KeyError。
  3. 编码与文件管理:添加encoding='utf-8'确保中文/特殊字符正常写入,with语句自动关闭文件。

内容的提问来源于stack exchange,提问作者dhrice

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:08:31