You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python高效从大型JSON文件中查找特定键值对并生成符合条件的URL列表CSV文件

针对你的需求,我整理了处理大型JSON文件的高效解决方案,核心思路是流式处理,避免把整个文件加载到内存导致溢出,这是处理大文件的关键:

一、高效提取以https://api开头的URL并生成CSV

处理大型JSON时,绝对不能用json.load()一次性加载整个文件,推荐用ijson库做流式解析,搭配csv模块逐行写入结果,全程内存占用极低。

步骤1:安装依赖库

pip install ijson

步骤2:实现代码

import ijson
import csv

def extract_api_urls(json_input_path, csv_output_path):
    # 同时打开源JSON和目标CSV文件,用上下文管理器自动处理关闭
    with open(json_input_path, 'r', encoding='utf-8') as json_file, \
         open(csv_output_path, 'w', newline='', encoding='utf-8') as csv_file:
        
        csv_writer = csv.writer(csv_file)
        csv_writer.writerow(['url'])  # 写入CSV表头
        
        # 流式解析JSON中所有"url"字段的值
        # 注意:`'item.url'`中的路径需要根据你的JSON结构调整
        # 如果你的JSON是顶层数组(比如[{"url": "...", ...}, {...}]),用`'item.url'`
        # 如果是嵌套结构,比如{"data": [{"url": "..."}]},则用`'data.item.url'`
        for url in ijson.items(json_file, 'item.url'):
            if url.startswith('https://api'):
                csv_writer.writerow([url])

# 调用示例
extract_api_urls('large_data.json', 'filtered_api_urls.csv')

为什么这是最高效的?

  • 流式解析:ijson会逐块读取JSON文件,只解析当前需要的"url"字段,不会把整个文件加载到内存,哪怕是几十GB的文件也能处理。
  • 逐行写入CSV:找到符合条件的URL就立即写入,不会把所有结果存在列表里占用内存。

如果你的JSON是JSON Lines格式(每行一个独立的JSON对象),可以用更轻量的方式,不需要安装ijson:

import json
import csv

def process_json_lines(jsonl_path, csv_output_path):
    with open(jsonl_path, 'r', encoding='utf-8') as jsonl_file, \
         open(csv_output_path, 'w', newline='', encoding='utf-8') as csv_file:
        
        csv_writer = csv.writer(csv_file)
        csv_writer.writerow(['url'])
        
        for line in jsonl_file:
            try:
                data = json.loads(line.strip())
                if 'url' in data and data['url'].startswith('https://api'):
                    csv_writer.writerow([data['url']])
            except json.JSONDecodeError:
                print(f"跳过无效行:{line[:50]}...")
                continue
二、Python从大型JSON文件中查找特定键值对

同样用流式解析的思路,避免内存过载,这里分两种场景:

场景1:查找所有包含指定键的项

比如查找所有带有"url"键的对象:

import ijson

def find_items_with_key(json_path, target_key):
    """用生成器返回结果,避免内存占用"""
    with open(json_path, 'r', encoding='utf-8') as json_file:
        # 还是要根据JSON结构调整路径,比如顶层数组用'item',嵌套数组用'data.item'
        for item in ijson.items(json_file, 'item'):
            if target_key in item:
                yield item

# 使用示例:遍历所有带"url"键的项
for item in find_items_with_key('large_data.json', 'url'):
    if item['url'].startswith('https://api'):
        print(item)

场景2:查找键值完全匹配的项

比如查找"status"键等于"success"的所有项:

import ijson

def find_items_by_key_value(json_path, target_key, target_value):
    with open(json_path, 'r', encoding='utf-8') as json_file:
        for item in ijson.items(json_file, 'item'):
            if item.get(target_key) == target_value:
                yield item

# 使用示例
for success_item in find_items_by_key_value('large_data.json', 'status', 'success'):
    print(success_item)

关键注意点

  • 路径适配:ijson的路径参数(比如'item.url')需要和你的JSON结构对应,你可以先用ijson.parse()查看JSON的结构,或者根据文件格式调整。
  • 生成器优先:用yield返回结果,而不是把所有结果存在列表里,这样内存始终保持低占用。

内容的提问来源于stack exchange,提问作者Dcook

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 00:52:34