You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用subprocess并行Python脚本时JSON解码报错求助

问题解决思路及方案

常见报错原因

  • 子进程输出包含非JSON内容(比如调试打印信息、爬虫日志、异常栈),导致json.loads无法解析
  • 子进程stdout捕获不完整,或编码格式不匹配导致输出乱码
  • 爬虫脚本未正确将字典转为标准JSON字符串输出(比如直接打印字典对象而非用json.dumps转换)

具体修复步骤

1. 修正爬虫脚本的输出逻辑

确保Magalu_Telefonia.py和Magalu_Eletrodom.py仅输出纯JSON字符串,删除所有调试用print语句:

# 爬虫脚本末尾添加以下代码(替换原有输出逻辑)
import json

# 假设爬取结果存储在result_dict变量中
print(json.dumps(result_dict, ensure_ascii=False))
  • 若脚本有异常处理,将异常信息输出到stderr而非stdout,避免干扰JSON解析
  • 单独运行脚本检查输出:控制台应只显示一段标准JSON文本,无其他内容

2. 优化主脚本的子进程调用

调整subprocess参数,确保正确捕获输出并处理编码:

import subprocess
import json
from concurrent.futures import ThreadPoolExecutor

def execute_spider(script_path):
    proc = subprocess.run(
        ["python", script_path],
        capture_output=True,
        text=True,
        encoding="utf-8",
        timeout=30
    )
    
    # 先检查脚本执行状态,输出错误信息用于排查
    if proc.returncode != 0:
        print(f"脚本{script_path}执行失败: {proc.stderr}")
        return None
    
    try:
        # 去除输出首尾空白字符后解析JSON
        return json.loads(proc.stdout.strip())
    except json.JSONDecodeError as e:
        print(f"解析{script_path}输出失败: {str(e)}")
        print(f"原始输出内容: {repr(proc.stdout)}") # 打印原始输出定位问题
        return None

# 并行执行两个爬虫
with ThreadPoolExecutor(max_workers=2) as executor:
    future_list = [
        executor.submit(execute_spider, "Magalu_Telefonia.py"),
        executor.submit(execute_spider, "Magalu_Eletrodom.py")
    ]
    final_results = [f.result() for f in future_list]

print("爬取结果:", final_results)

3. 排查验证要点

  • 单独运行每个爬虫脚本,确认输出仅为标准JSON格式
  • 若输出包含特殊字符,确保json.dumps添加ensure_ascii=False参数
  • 若仍报错,查看主脚本打印的原始输出内容,定位混入的非JSON文本并从爬虫脚本中移除

内容的提问来源于stack exchange,提问作者Davi Riani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 17:07:14