You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python食谱管理项目:动态写入JSON文件的技术咨询

Python食谱管理工具:JSON动态写入实现方案

Python内置的json库就能满足你的JSON文件动态写入需求,无需额外安装第三方库,完美适配你用requests和BeautifulSoup的爬取流程,具体实现逻辑如下:

核心思路

动态写入JSON的关键是先读取已有数据(若文件存在),再追加新爬取的URL,最后重新写入文件,同时处理文件不存在、重复URL等常见场景。

代码示例

import json
import requests
from bs4 import BeautifulSoup

# 爬取食谱URL的函数(替换成你的实际爬取逻辑)
def crawl_recipe_urls(keyword):
    # 这里编写你的requests请求+BeautifulSoup解析代码
    # 示例返回模拟的URL列表
    return [
        f"https://example.com/recipe/{keyword}-001",
        f"https://example.com/recipe/{keyword}-002"
    ]

# 动态更新JSON文件的函数
def update_recipe_json(new_urls, json_path="recipes.json"):
    # 读取已有数据或初始化空列表
    try:
        with open(json_path, "r", encoding="utf-8") as f:
            existing_data = json.load(f)
            # 确保数据是列表格式,兼容异常文件内容
            if not isinstance(existing_data, list):
                existing_data = []
    except FileNotFoundError:
        existing_data = []
    
    # 去重:避免重复写入相同URL
    unique_new_urls = [url for url in new_urls if url not in existing_data]
    existing_data.extend(unique_new_urls)
    
    # 写入JSON文件,格式化输出更易读
    with open(json_path, "w", encoding="utf-8") as f:
        json.dump(existing_data, f, ensure_ascii=False, indent=4)

# 调用示例
if __name__ == "__main__":
    target_keyword = "番茄炒蛋"
    crawled_urls = crawl_recipe_urls(target_keyword)
    update_recipe_json(crawled_urls)

关键细节说明

  • json.load():将JSON文件内容转换为Python列表,方便后续追加操作
  • json.dump():将处理后的Python列表写入JSON文件,ensure_ascii=False支持中文正常显示,indent=4让JSON结构更清晰易读
  • 去重逻辑:过滤掉已存在的URL,避免数据冗余
  • 异常处理:当JSON文件不存在时自动初始化空列表,避免程序报错

内容的提问来源于stack exchange,提问作者Ruixuan G

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.01 13:25:00