You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取图片报错:no such element 问题排查求助

问题:爬取自有网站帖子时遇到「no such element」错误

爬取自有网站帖子时,需要提取帖子里的图片和文本,但触发了「no such element」错误。明明目标帖子里有src元素和文本,找不到报错原因,求技术建议。

现有代码

import csv
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
# disable chromedriver log message in cmd
options.add_experimental_option("excludeSwitches", ["enable-automation", "enable-logging"])

service = Service(executable_path="C:\webdrivers\chromedriver.exe")
driver = webdriver.Chrome(service=service, options=options)
wait = WebDriverWait(driver, 10)

driver.get("https://navalcommand.enjin.com/forum/m/11178354/viewforum/2989688/page/1")
# get number of pages
num_pages = driver.find_element(By.CSS_SELECTOR, "span.text.rightmost").text.split(' ')[1]

for page in range(2, int(num_pages)):
    # find all threads on the current page
    threads = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "a.thread-view.thread-subject")))
    # get links to threads
    thread_links = [x.get_attribute('href') for x in threads]
    # open each link and get all the posts in thread
    for link in thread_links:
        driver.get(link)
        thread_content = driver.find_elements(By.CSS_SELECTOR, "div.post-content")
        # get thread id
        thread_id = driver.current_url.split('d/')[1].split('-')[0]
        # save received data in csv
        for post in thread_content:
            post_content = post.text or post.find_element(By.TAG_NAME, 'img').get_attribute('src')
            with open(file=r'C:\Users\jammi\OneDrive\Desktop\Navcom\aar\ ' f'{thread_id}_naval-command.csv', mode='a',
                      encoding="utf-8") as f:
                writer = csv.writer(f, lineterminator='\n')
                writer.writerow([post_content])
    driver.get(f"https://navalcommand.enjin.com/forum/viewforum/2989694/m/11178354/page/{page}")

driver.quit()

页面截图

页面截图


技术解决方案

错误根源

  1. 内容提取逻辑缺陷:post.text or post.find_element(...)逻辑不严谨,若帖子无图片元素,find_element会直接抛出NoSuchElementException;
  2. 缺少元素等待:进入帖子页面后直接提取内容,可能页面未加载完成导致找不到元素;
  3. 页码循环范围错误:range(2, int(num_pages))会漏掉最后一页(range为左闭右开);
  4. 文件操作效率低下:每次写入都打开文件,易引发IO问题。

修复步骤

  1. 增加帖子内容加载等待
    替换原代码中thread_content = driver.find_elements(...)为:

    thread_content = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.post-content")))
    
  2. 修改内容提取逻辑,兼容多场景
    用find_elements(复数)替代find_element,避免找不到元素时抛出异常,同时分别处理文本和图片:

    for post in thread_content:
        text_content = post.text.strip()
        # 查找所有img元素,找不到返回空列表
        img_elements = post.find_elements(By.TAG_NAME, 'img')
        
        if text_content:
            post_content = text_content
        elif img_elements:
            # 取第一个图片的src
            post_content = img_elements[0].get_attribute('src')
        else:
            post_content = ""  # 空内容的情况可按需处理
    
  3. 修正页码循环范围
    将range(2, int(num_pages))改为:

    for page in range(2, int(num_pages) + 1):
    
  4. 优化CSV写入逻辑
    把文件打开操作移到单条帖子循环外,减少IO次数:

    for link in thread_links:
        driver.get(link)
        thread_content = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.post-content")))
        thread_id = driver.current_url.split('d/')[1].split('-')[0]
        csv_path = rf'C:\Users\jammi\OneDrive\Desktop\Navcom\aar\{thread_id}_naval-command.csv'
        
        with open(csv_path, mode='a', encoding="utf-8", newline='') as f:
            writer = csv.writer(f)
            for post in thread_content:
                text_content = post.text.strip()
                img_elements = post.find_elements(By.TAG_NAME, 'img')
                
                if text_content:
                    post_content = text_content
                elif img_elements:
                    post_content = img_elements[0].get_attribute('src')
                else:
                    post_content = ""
                writer.writerow([post_content])
    

    注:newline=''可避免Windows下CSV出现空行。

  5. 增强页码获取的健壮性
    替换原页码提取代码,避免因页面文本格式变化导致索引错误:

    page_info_text = driver.find_element(By.CSS_SELECTOR, "span.text.rightmost").text
    num_pages = page_info_text.split()[-1]  # 取文本最后一个数字,适配"Page 1 of N"格式
    

内容的提问来源于stack exchange,提问作者Jherky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 05:10:21