You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

读取JSON文件批量解析网站时出现JSONDecodeError求助

解决JSONDecodeError问题的方案

错误原因

你的JSON文件是完整的JSON数组结构,但代码中逐行读取并尝试解析每一行。例如第一行的[、带逗号的行、空行,这些单独的行都不是有效的JSON对象,因此触发JSONDecodeError: Expecting value。

解决方案

方案一:一次性加载整个JSON文件(推荐)

直接读取整个JSON文件解析为数组,再遍历数组中的每个对象,这是最符合JSON规范的做法,代码简洁且可靠:

import time
import json

# 一次性加载整个JSON数组
with open('Codes_test.json', 'r', encoding='UTF-8') as content:
    url_items = json.load(content)

counter = 0
for item in url_items:
    counter += 1
    if counter > 5:
        break
    
    url = item["page_url"]
    print(f'Retrieving data for {url}')
    
    retrieved_data = parse_website(url)
    
    # 使用with语句自动管理文件,避免手动关闭遗漏
    with open('page_content.json', 'a', encoding='utf-8') as f:
        f.write(json.dumps(retrieved_data))
        f.write('\n')
    
    time.sleep(1)

方案二:逐行读取并过滤无效行(仅适用于超大文件)

如果JSON文件过大无法一次性加载,可以过滤掉数组符号、逗号、空行,只解析有效的对象行:

import time
import json

with open('Codes_test.json', 'r', encoding='UTF-8') as content:
    counter = 0
    for line in content:
        cleaned_line = line.strip()
        # 跳过非对象行
        if cleaned_line in ('[', ']', '') or cleaned_line.endswith(','):
            continue
        
        counter += 1
        if counter > 5:
            break
        
        # 去掉行尾逗号后解析JSON对象
        obj = json.loads(cleaned_line.rstrip(','))
        url = obj["page_url"]
        print(f'Retrieving data for {url}')
        
        retrieved_data = parse_website(url)
        
        with open('page_content.json', 'a', encoding='utf-8') as f:
            f.write(json.dumps(retrieved_data))
            f.write('\n')
        
        time.sleep(1)

额外优化建议

  • 始终使用with语句操作文件,它会自动处理文件关闭,避免资源泄漏。
  • 循环中保留time.sleep(1)可以降低对目标网站的请求频率,减少被反爬机制拦截的概率。

内容的提问来源于stack exchange,提问作者Bjorn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 21:45:29