You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Selenium爬取的数据存入JSON文件且避免title字段重复

问题根源
  • 每次调用函数时都新建了仅包含当前章节的书籍对象,直接拼接写入文件导致title重复
  • JSON是结构化格式,不能直接追加字符串到文件末尾,必须读取全量内容修改后再整体写入
  • 使用了Python内置关键字list作为变量名,存在语法隐患
  • 每次循环都新建Chrome浏览器实例,资源消耗高、运行效率低
修复方案

调整读写逻辑:每次写入前先读取已有JSON文件内容,判断目标书籍是否已存在,存在就追加章节到对应chapters列表,不存在就新增书籍条目,最后将完整结构统一序列化写回文件。

完整修改代码
import json
from selenium import webdriver
from selenium.webdriver.chrome.options import Options

def getAllImages(driver, url, book_title="Kingdom"):
    try:
        driver.get(url)
        driver.implicitly_wait(2)
    except Exception as e:
        print("Error Getting Images Page :", e)
        return
    
    current_chapter_title = driver.title
    divs = driver.find_elements_by_class_name("page-break ")
    images = []
    for div in divs:
        imgs = div.find_elements_by_tag_name("img")
        for img in imgs:
            images.append(img.get_attribute("src").strip())
    
    # 构造当前章节对象
    chapter = {
        "chapter-title": current_chapter_title,
        "images": images
    }

    # 读取已有JSON数据
    file_path = "./manga.json"
    book_list = []
    try:
        with open(file_path, "r", encoding="utf-8") as f:
            book_list = json.load(f)
    except (FileNotFoundError, json.JSONDecodeError):
        # 文件不存在或者内容为空/格式错误,初始化空列表
        book_list = []
    
    # 查找目标书籍是否已存在
    target_book = None
    for book in book_list:
        if book.get("title") == book_title:
            target_book = book
            break
    
    if target_book:
        # 存在就追加章节
        target_book["chapters"].append(chapter)
    else:
        # 不存在就新增书籍条目
        new_book = {
            "title": book_title,
            "chapters": [chapter]
        }
        book_list.append(new_book)
    
    # 整体写回文件
    with open(file_path, "w", encoding="utf-8") as f:
        json.dump(book_list, f, ensure_ascii=False, indent=2)
    
    print(f"章节 {current_chapter_title} 写入成功 🎉")

if __name__ == "__main__":
    # 提前初始化driver,避免每次循环新建实例,提升效率
    chrome_options = Options()
    chrome_options.add_argument("--headless")
    driver = webdriver.Chrome(options=chrome_options)
    
    links = [] # 填入你的章节链接列表
    for url in links:
        # 后续爬取其他书时,修改book_title参数即可自动生成新的书籍条目
        getAllImages(driver, url, book_title="Kingdom")
    
    driver.quit()
注意事项
  • 最终输出的是标准JSON数组结构,符合API数据源的使用要求,直接解析即可获取所有书籍数据
  • 代码兼容第一次运行时文件不存在的场景,无需提前手动创建manga.json
  • 不要手动修改manga.json内容,避免JSON格式错误导致读取失败

内容的提问来源于stack exchange,提问作者ABD JALIL Danane

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 21:24:04