如何将icrawler GoogleImageCrawler下载的图片存入SQLite?
解决GoogleImageCrawler下载图片存入SQLite的问题
你的代码存在几个关键问题导致无法将图片存入数据库:
get_photo函数没有返回图片的二进制数据,插入时第二个参数为None,违反了catphoto BLOB NOT NULL的约束- 循环中两次调用
get_photo,会重复下载相同图片,浪费资源 - 无法定位下载后的具体图片文件,无法读取其二进制内容
修改方案
- 让
get_photo函数完成下载后,读取图片的二进制数据并返回 - 循环中仅调用一次
get_photo,避免重复下载 - 添加异常处理,应对下载失败或插入数据库失败的情况
完整代码
from icrawler.builtin import GoogleImageCrawler import sqlite3 import os # 初始化数据库连接 database = sqlite3.connect("parserwoman.db") sql = database.cursor() table_name = 'photos' # 创建表(优化语句) sql.execute(f"CREATE TABLE IF NOT EXISTS {table_name} (title TEXT PRIMARY KEY NOT NULL, catphoto BLOB NOT NULL)") def get_photo(keywords: str): filters = dict(size='large', type='photo') download_dir = r"C:\Users\Acer\Desktop\photo" # 确保下载目录存在 os.makedirs(download_dir, exist_ok=True) google_crawler = GoogleImageCrawler(storage={'root_dir': download_dir}) google_crawler.crawl( keyword=keywords, max_num=1, min_size=(1024, 1024), max_size=None, filters=filters, file_idx_offset='auto' ) # 筛选下载目录中的图片文件 image_extensions = ('.jpg', '.jpeg', '.png', '.gif') image_files = [ f for f in os.listdir(download_dir) if f.lower().endswith(image_extensions) ] if not image_files: raise FileNotFoundError(f"未找到关键词 {keywords} 对应的下载图片") # 按文件修改时间排序,取最新下载的图片 image_files.sort( key=lambda x: os.path.getmtime(os.path.join(download_dir, x)), reverse=True ) latest_image_path = os.path.join(download_dir, image_files[0]) # 读取图片二进制数据 with open(latest_image_path, 'rb') as img_file: image_data = img_file.read() # 可选:如果不需要保留本地图片,下载后删除 # os.remove(latest_image_path) return image_data titles = [ 'Abyssinian', 'American Bobtail', 'Himalayan' ] # 循环插入数据 for element in titles: try: img_data = get_photo(element) sql.execute( f"INSERT INTO {table_name} (title, catphoto) VALUES (?, ?)", (element, img_data) ) database.commit() print(f"{element} 图片已成功存入数据库") except sqlite3.IntegrityError: print(f"{element} 已存在于数据库中,跳过插入") except Exception as e: print(f"处理 {element} 时出错: {str(e)}") database.rollback() # 关闭数据库连接 database.close()
关键说明
- 函数会自动筛选下载目录中的图片文件,通过修改时间找到最新下载的那张(因为
max_num=1,所以只会有一个符合的文件) - 添加了主键重复的异常处理,避免重复插入相同标题的数据
- 可选删除本地图片的代码,如需节省本地空间可以取消注释
内容的提问来源于stack exchange,提问作者Atom
相关产品推荐
相关产品推荐

