You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Rails中用Selenium等库下载网站图片遇错,求解决

解决Ruby中Selenium+Nokogiri抓取图片并下载到Ubuntu的问题

咱们先逐个拆解你遇到的每个尝试的问题,然后给出靠谱的实现方式:


先修正基础代码里的小笔误

首先你的基础代码有两处小问题:一是require 'open-url'应该是require 'open-uri';二是parsed_content = Nokogiri::HTML(content)里的content要改成你之前定义的document,不然会直接报错。修正后的基础代码更严谨:

require 'open-uri'
require 'selenium-webdriver'
require 'nokogiri'
require 'uri'
require 'fileutils' # 后续创建目录需要用到这个库

url = "https://www.google.com/"
browser = Selenium::WebDriver.for :chrome
document = open(url).read
parsed_content = Nokogiri::HTML(document)
image_src = parsed_content.css('#hplogo').attr('src').value
image_url = URI.join(url, image_src).to_s # 用URI.join自动处理相对路径,比手动拼接更安全
binding.pry

逐个分析你的尝试错误

1) 第一次尝试的问题

你用~/image.png作为保存路径,但Ruby的文件操作不会自动解析波浪号(~),它会把~当成普通字符处理。另外你创建目录的代码也错了:FileUtils.mkdir_p(image)里的image是图片URL,不是目录路径,完全不对。

修正方法:用File.expand_path解析波浪号,先确保保存目录存在:

save_path = File.expand_path('~/image.png')
# 先创建保存路径的父目录(如果不存在)
FileUtils.mkdir_p(File.dirname(save_path)) unless File.directory?(File.dirname(save_path))

download = open(image_url)
IO.copy_stream(download, save_path)

2) 第二次尝试的问题

返回#<File:image.png (closed)>说明文件已经被成功写入并关闭了!你找不到文件是因为不知道当前工作目录在哪里。在pry里输入Dir.pwd就能看到脚本运行的目录,image.png就保存在那里。

如果想指定到你能快速找到的路径(比如桌面),直接用绝对路径:

save_path = File.expand_path('~/Desktop/image.png')
open(save_path, 'wb') do |file|
  file << open(image_url).read
end

这里的wb模式很重要,因为图片是二进制文件,用二进制写入模式能避免编码乱码问题。

3) 第三次尝试的问题

返回的5482是写入文件的字节数,说明文件已经成功保存了!同样是路径的问题,用Dir.pwd查看当前目录,就能找到image.png。如果要指定路径,参考上面的绝对路径写法:

save_path = File.expand_path('~/Downloads/image.png')
IO.copy_stream(open(image_url), save_path)

4) 第四次尝试的问题

你把File.open的参数顺序搞反了!File.open的第一个参数是文件名,第二个参数是打开模式,你写成了File.open(image, 'image.png'),把图片URL当成了文件名,把文件名当成了模式,自然会报错invalid access mode image.png。

正确的写法应该是:

save_path = File.expand_path('~/image.png')
read_image = open(image_url).read
File.open(save_path, 'wb') do |file|
  file.write(read_image) # 用write比puts更适合二进制内容,puts会自动加换行符破坏图片
end

最推荐的实现方式

综合来看,用IO.copy_stream配合绝对路径是最简洁高效的,因为它不需要把整个图片内容读到内存里,适合大文件:

require 'open-uri'
require 'nokogiri'
require 'fileutils'

url = "https://www.google.com/"
document = open(url).read
parsed_content = Nokogiri::HTML(document)
image_src = parsed_content.css('#hplogo').attr('src').value
image_url = URI.join(url, image_src).to_s

save_path = File.expand_path('~/Downloads/google_logo.png')
FileUtils.mkdir_p(File.dirname(save_path))

IO.copy_stream(open(image_url), save_path)
puts "图片已保存到:#{save_path}"

运行后,你就能在~/Downloads目录下找到google_logo.png了。

内容的提问来源于stack exchange,提问作者Joe Morano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:29:39