Python食谱管理项目:动态写入JSON文件的技术咨询
Python食谱管理工具:JSON动态写入实现方案
Python内置的json库就能满足你的JSON文件动态写入需求,无需额外安装第三方库,完美适配你用requests和BeautifulSoup的爬取流程,具体实现逻辑如下:
核心思路
动态写入JSON的关键是先读取已有数据(若文件存在),再追加新爬取的URL,最后重新写入文件,同时处理文件不存在、重复URL等常见场景。
代码示例
import json import requests from bs4 import BeautifulSoup # 爬取食谱URL的函数(替换成你的实际爬取逻辑) def crawl_recipe_urls(keyword): # 这里编写你的requests请求+BeautifulSoup解析代码 # 示例返回模拟的URL列表 return [ f"https://example.com/recipe/{keyword}-001", f"https://example.com/recipe/{keyword}-002" ] # 动态更新JSON文件的函数 def update_recipe_json(new_urls, json_path="recipes.json"): # 读取已有数据或初始化空列表 try: with open(json_path, "r", encoding="utf-8") as f: existing_data = json.load(f) # 确保数据是列表格式,兼容异常文件内容 if not isinstance(existing_data, list): existing_data = [] except FileNotFoundError: existing_data = [] # 去重:避免重复写入相同URL unique_new_urls = [url for url in new_urls if url not in existing_data] existing_data.extend(unique_new_urls) # 写入JSON文件,格式化输出更易读 with open(json_path, "w", encoding="utf-8") as f: json.dump(existing_data, f, ensure_ascii=False, indent=4) # 调用示例 if __name__ == "__main__": target_keyword = "番茄炒蛋" crawled_urls = crawl_recipe_urls(target_keyword) update_recipe_json(crawled_urls)
关键细节说明
json.load():将JSON文件内容转换为Python列表,方便后续追加操作json.dump():将处理后的Python列表写入JSON文件,ensure_ascii=False支持中文正常显示,indent=4让JSON结构更清晰易读- 去重逻辑:过滤掉已存在的URL,避免数据冗余
- 异常处理:当JSON文件不存在时自动初始化空列表,避免程序报错
内容的提问来源于stack exchange,提问作者Ruixuan G
相关产品推荐
相关产品推荐

