You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取小说章节内容遇阻,求可行解决方案

问题排查与解决方案

核心问题定位

  • 直接用requests发送无请求头的请求时,目标站点会触发反爬机制,返回的HTML并非浏览器中看到的正常渲染页面,自然找不到id=chapterContent的标签
  • 原代码存在基础语法错误:Python布尔真值为大写True,写为true会触发未定义变量报错
  • 原种子URL没有占位符,无法适配多页迭代爬取逻辑,需要调整为带参数占位符的格式

修复后可运行代码

import requests
import re
import time
import os
from bs4 import BeautifulSoup

def browse_and_scrape(seed_url, page_id, save_path="./novel_chapters/"):
    # 新建章节存储文件夹
    if not os.path.exists(save_path):
        os.makedirs(save_path)
    # 构造浏览器请求头绕过基础反爬
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"
    }
    # 拼接当前章节URL
    formatted_url = seed_url.format(page_id)
    try:
        html_text = requests.get(formatted_url, headers=headers).text
        soup = BeautifulSoup(html_text, "html.parser")
        # 定位章节内容标签
        content_tag = soup.find(id="chapterContent")
        if not content_tag:
            return f"未找到章节内容,URL: {formatted_url}"
        # 提取纯文本内容
        chapter_content = content_tag.get_text(strip=True)
        # 提取章节名作为存储文件名
        chapter_title = soup.find("h1", class_="chapter-title").get_text(strip=True)
        # 写入本地文本文件
        with open(f"{save_path}/{chapter_title}.txt", "w", encoding="utf-8") as f:
            f.write(chapter_content)
        print(f"章节《{chapter_title}》爬取存储完成")
        # 提取下一页ID实现自动迭代
        next_chapter_tag = soup.find("a", id="nextChapter")
        if next_chapter_tag and next_chapter_tag.get("href"):
            next_href = next_chapter_tag["href"]
            next_id = re.search(r"chapter/(\d+)-", next_href).group(1)
            return next_id
        return None
    except Exception as e:
        return f"请求出错:{str(e)}"

if __name__ == "__main__":
    # 带占位符的种子URL
    seed_url = "http://wnmtl.org/chapter/{}-temp.html"
    # 起始章节ID,当前测试章节为324909
    current_id = "324909"
    print("爬取开始")
    # 循环爬取直到无下一页
    while current_id:
        res = browse_and_scrape(seed_url, current_id)
        if isinstance(res, str) and res.isdigit():
            current_id = res
            # 加2秒延时避免请求频率过高被封禁IP
            time.sleep(2)
        else:
            if res:
                print(f"爬取中断:{res}")
            break
    print("爬取流程结束")

额外注意事项

  • 存储时使用utf-8编码可避免中文/特殊字符乱码
  • 如果后续站点反爬规则升级,可以换用selenium模拟完整浏览器加载页面后再提取内容

内容的提问来源于stack exchange,提问作者uche ozor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 06:06:07