使用BeautifulSoup爬取小说章节内容遇阻,求可行解决方案
问题排查与解决方案
核心问题定位
- 直接用
requests发送无请求头的请求时,目标站点会触发反爬机制,返回的HTML并非浏览器中看到的正常渲染页面,自然找不到id=chapterContent的标签 - 原代码存在基础语法错误:Python布尔真值为大写
True,写为true会触发未定义变量报错 - 原种子URL没有占位符,无法适配多页迭代爬取逻辑,需要调整为带参数占位符的格式
修复后可运行代码
import requests import re import time import os from bs4 import BeautifulSoup def browse_and_scrape(seed_url, page_id, save_path="./novel_chapters/"): # 新建章节存储文件夹 if not os.path.exists(save_path): os.makedirs(save_path) # 构造浏览器请求头绕过基础反爬 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36" } # 拼接当前章节URL formatted_url = seed_url.format(page_id) try: html_text = requests.get(formatted_url, headers=headers).text soup = BeautifulSoup(html_text, "html.parser") # 定位章节内容标签 content_tag = soup.find(id="chapterContent") if not content_tag: return f"未找到章节内容,URL: {formatted_url}" # 提取纯文本内容 chapter_content = content_tag.get_text(strip=True) # 提取章节名作为存储文件名 chapter_title = soup.find("h1", class_="chapter-title").get_text(strip=True) # 写入本地文本文件 with open(f"{save_path}/{chapter_title}.txt", "w", encoding="utf-8") as f: f.write(chapter_content) print(f"章节《{chapter_title}》爬取存储完成") # 提取下一页ID实现自动迭代 next_chapter_tag = soup.find("a", id="nextChapter") if next_chapter_tag and next_chapter_tag.get("href"): next_href = next_chapter_tag["href"] next_id = re.search(r"chapter/(\d+)-", next_href).group(1) return next_id return None except Exception as e: return f"请求出错:{str(e)}" if __name__ == "__main__": # 带占位符的种子URL seed_url = "http://wnmtl.org/chapter/{}-temp.html" # 起始章节ID,当前测试章节为324909 current_id = "324909" print("爬取开始") # 循环爬取直到无下一页 while current_id: res = browse_and_scrape(seed_url, current_id) if isinstance(res, str) and res.isdigit(): current_id = res # 加2秒延时避免请求频率过高被封禁IP time.sleep(2) else: if res: print(f"爬取中断:{res}") break print("爬取流程结束")
额外注意事项
- 存储时使用
utf-8编码可避免中文/特殊字符乱码 - 如果后续站点反爬规则升级,可以换用
selenium模拟完整浏览器加载页面后再提取内容
内容的提问来源于stack exchange,提问作者uche ozor
相关产品推荐
相关产品推荐

