You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用icrawler时如何获取图片的源URL?

获取icrawler中图片的源URL方法

你可以通过自定义icrawler的组件,直接在代码中捕获并保存图片的源URL,无需依赖终端输出,以下是几种实用方法:

方法一:自定义下载器捕获URL

重写ImageDownloader类,在处理图片下载前提取并保存源URL:

from icrawler.builtin import GoogleImageCrawler
from icrawler.downloader import ImageDownloader
from six.moves.urllib.parse import urlparse

class CustomDownloader(ImageDownloader):
    def get_filename(self, task, default_ext):
        # 提取图片源URL
        img_url = task['file_url']
        # 打印URL或保存到文件
        print(f"图片源URL:{img_url}")
        with open('image_urls.txt', 'a', encoding='utf-8') as f:
            f.write(img_url + '\n')
        # 保留原命名逻辑(也可自定义文件名)
        return super().get_filename(task, default_ext)

# 初始化爬虫并使用自定义下载器
crawler = GoogleImageCrawler(
    downloader_cls=CustomDownloader,
    storage={'root_dir': 'your_image_save_dir'}
)
crawler.crawl(keyword='programing gif', max_num=10)

方法二:自定义解析器提取URL

在解析搜索结果页面时直接获取图片URL:

from icrawler.builtin import GoogleImageCrawler
from icrawler.parser import GoogleParser

class CustomParser(GoogleParser):
    def parse(self, response):
        # 调用父类方法解析出图片任务列表
        tasks = super().parse(response)
        # 遍历任务提取URL
        for task in tasks:
            img_url = task['file_url']
            print(f"解析到图片URL:{img_url}")
            with open('image_urls.txt', 'a', encoding='utf-8') as f:
                f.write(img_url + '\n')
        return tasks

# 使用自定义解析器初始化爬虫
crawler = GoogleImageCrawler(
    parser_cls=CustomParser,
    storage={'root_dir': 'your_image_save_dir'}
)
crawler.crawl(keyword='programing gif', max_num=10)

方法三:利用信号机制绑定回调(部分版本支持)

如果你的icrawler版本支持信号,可以绑定before_download信号来获取URL:

from icrawler.builtin import GoogleImageCrawler

def capture_image_url(sender, task):
    img_url = task['file_url']
    print(f"即将下载图片URL:{img_url}")
    with open('image_urls.txt', 'a', encoding='utf-8') as f:
        f.write(img_url + '\n')

crawler = GoogleImageCrawler(storage={'root_dir': 'your_image_save_dir'})
# 绑定回调函数到下载前信号
crawler.downloader.signals.connect(capture_image_url, signal='before_download')

crawler.crawl(keyword='programing gif', max_num=10)

以上方法都会将图片源URL保存到image_urls.txt文件中,同时也会在控制台打印,方便你批量获取和管理这些URL。

内容的提问来源于stack exchange,提问作者mrithul e

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 00:12:51