爬取鱼类图片:提取img的src并生成分组嵌套列表问题求助
解决鱼类图片按鱼种分组的嵌套列表爬取问题
你的核心问题是没有为每个鱼种创建独立的子列表存储图片src,当前代码直接将所有src追加到全局列表,导致结构扁平。以下是修正后的代码及关键说明:
修正后的代码
from selenium import webdriver from selenium.common.exceptions import NoSuchElementException from bs4 import BeautifulSoup import time from selenium.webdriver.common.by import By species_with_foto = ['/fangster/aborre-perca-fluviatilis/1', '/fangster/almindelig-tangnaal-syngnathus-typhle/155', '/fangster/ansjos-engraulis-encrasicholus/66', '/fangster/atlantisk-tun-blaafinnet-tun-thunnus-thynnus-/137'] titles = [] fish_photos = [] driver = webdriver.Chrome() # 根据你的浏览器类型调整驱动 for x in species_with_foto: specie_page = 'https://www.fiskefoto.dk' + x driver.get(specie_page) soup = BeautifulSoup(driver.page_source, 'html.parser') titles.append(x) # 为当前鱼种创建临时列表,专门存储该鱼种的图片src current_species_photos = [] images = soup.find_all('img', attrs={'class':'rapportBillede'}) for img in images: if img.has_attr('src'): current_species_photos.append(img['src']) # 将当前鱼种的图片列表加入全局嵌套列表 fish_photos.append(current_species_photos) try: driver.find_element(by=By.XPATH, value='/html/body/form/div[4]/div[1]/div/div[13]/div[2]/div/div').click() print('Clicked next', x) except NoSuchElementException: print('Successfully finished -', x) time.sleep(2) driver.quit() # 验证输出结构 print(fish_photos)
关键修改点
- 统一使用selenium driver获取页面内容(原代码混合
urlopen和driver会导致页面状态不一致,因为点击下一页操作依赖driver) - 每个鱼种循环内初始化
current_species_photos临时列表,单独存储当前鱼种的所有图片src - 将临时列表追加到
fish_photos,而非直接追加单个src,确保生成[[鱼种1图片列表], [鱼种2图片列表]]的嵌套结构 - 移除冗余的
KeyError捕获(已经通过has_attr('src')判断,不会触发键错误)
预期结果
修正后fish_photos会生成你需要的嵌套结构,示例如下:
[ ['/images/400/aborre-perca-fluviatilis-medefiskeri-bundrig-0,220kg-24cm-striber-rygfinne-regnorm-majs-spinner-358-22-29-14-740-2013-21-4.jpg', ...], ['/images/400/almindelig-tangnaal-syngnathus-typhle-...jpg', ...], # 后续鱼种的图片列表 ]
内容的提问来源于stack exchange,提问作者jmChrist
相关产品推荐
相关产品推荐

