使用epubjs提取EPUB图片遇404错误,求解决方案
EPUB.js提取EPUB图片返回404的解决方法
问题描述
我已成功使用epubjs读取EPUB文件,代码如下:
let book = ePub('epub/' + file);
随后尝试遍历文件提取文本和图片,使用的代码为:
book.loaded.spine.then(spine => { spine.each((item, ind) => { item.load(book.load.bind(book)).then(contents => { // console.log(contents.innerText) // console.log(contents); let body = contents.lastElementChild, hasimg = body.querySelector('img'); if (hasimg) { let img = new Image(); img.src = hasimg.src; //--------> 返回404 document.body.append(img); } }); }); });
文本提取正常,但图片无法获取:检测到hasimg包含正确的图片路径,不过设置img.src时返回404。请问如何将EPUB中的图片提取到网页中?是否需要先渲染每个“页面”才能获取图片?
问题原因
EPUB内部的图片路径是相对包内的路径,直接赋值给img.src时,浏览器会以当前网页的域名作为根路径去请求,自然找不到对应资源,导致404错误。无需提前渲染页面,利用epubjs自带的资源处理方法即可解决。
解决方法
方法1:用book.resolve()解析路径
使用epubjs提供的book.resolve()方法,将EPUB内部的相对路径转换为浏览器可识别的绝对URL:
if (hasimg) { let img = new Image(); // 解析图片的相对路径为绝对URL img.src = book.resolve(hasimg.getAttribute('src')); document.body.append(img); }
方法2:直接获取图片Blob数据(更稳定)
如果路径解析仍有问题,可以直接从EPUB包中加载图片为Blob,再转换为Data URL:
if (hasimg) { const imgRelativePath = hasimg.getAttribute('src'); // 加载图片资源并转为Blob book.load(imgRelativePath).then(response => response.blob()).then(blob => { let img = new Image(); img.src = URL.createObjectURL(blob); document.body.append(img); // 图片加载完成后释放URL,避免内存泄漏 img.onload = () => URL.revokeObjectURL(img.src); }); }
额外说明
不需要提前渲染页面就能提取图片,核心是借助epubjs的资源加载与解析能力,而非直接使用DOM元素的src属性(该属性对应的路径仅在EPUB内部有效,无法被浏览器直接识别)。
内容的提问来源于stack exchange,提问作者Andrew Foulds
相关产品推荐
相关产品推荐

