You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium/Python获取Noveltop.net章节下拉框的所有选项?

解决Noveltop.net章节下拉框加载问题的实操方案

1. 先排查iframe嵌套可能性

很多动态元素会被放在iframe里,先确认目标元素的位置:

  • 按F12打开浏览器开发者工具,定位那个空div,看它的父级有没有<iframe>标签
  • 如果存在iframe,必须先切换到对应iframe再操作,示例代码(Python):
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 等待iframe加载并切换(替换成实际的iframe ID/name/index)
iframe = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.ID, "target-iframe"))
)
driver.switch_to.frame(iframe)

# 尝试定位select元素
try:
    select = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.TAG_NAME, "select"))
    )
    # 获取所有选项
    options = select.find_elements(By.TAG_NAME, "option")
    for opt in options:
        print(opt.text, opt.get_attribute("value"))
finally:
    # 操作完成切回主页面
    driver.switch_to.default_content()

2. 触发懒加载(点击/滚动)

不少网站的下拉框是懒加载的,需要触发操作才会加载:

  • 先定位那个空div,点击它试试:
# 定位空div容器(替换成实际的选择器)
empty_div = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, "div.chapter-select-wrap"))
)
empty_div.click()

# 等待select加载完成
select = WebDriverWait(driver, 5).until(
    EC.presence_of_element_located((By.TAG_NAME, "select"))
)
  • 如果点击没用,试试滚动到该div的可视区域:
from selenium.webdriver.common.action_chains import ActionChains

ActionChains(driver).move_to_element(empty_div).perform()

3. 直接执行JS提取数据

如果元素是纯JS渲染的,绕开DOM直接拿数据更高效:

  • 打开浏览器控制台(F12),搜索chapter相关的全局变量(比如chapterList、chapters),确认是否有现成的章节数据
  • 用Selenium执行JS提取,示例:
# 执行JS获取全局章节数据(替换成实际变量名)
chapter_list = driver.execute_script("return window.chapterList;")
if chapter_list:
    for chap in chapter_list:
        print(chap.get("title"), chap.get("url"))
  • 也可以从页面的脚本标签里提取:
script_tags = driver.find_elements(By.TAG_NAME, "script")
for script in script_tags:
    content = script.get_attribute("innerHTML")
    if "chapterList" in content:
        # 用正则提取JSON格式的章节数据
        import re
        import json
        match = re.search(r'chapterList:\s*(\[.*?\])', content, re.DOTALL)
        if match:
            chapters = json.loads(match.group(1))
            for chap in chapters:
                print(chap.get("name"), chap.get("href"))

4. 检查自定义加载状态

有些网站不用document.readyState判断加载完成,需要等特定标识:

# 等待页面加载完成的标识元素出现(替换成实际选择器)
WebDriverWait(driver, 15).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, "div.page-loading.completed"))
)

新手必看提示

  • 永远先用浏览器F12手动观察:刷新页面,看select是何时出现的,有没有伴随点击/滚动操作
  • 定位器要准确:用F12的“复制选择器”功能获取元素的CSS/XPath,避免手写出错
  • 若遇到反爬,可给浏览器加配置(比如Chrome的--disable-blink-features=AutomationControlled),或者添加自定义请求头

内容的提问来源于stack exchange,提问作者user10463779

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 07:40:24