网页索引爬取及层级文件夹构建技术咨询
解决方案:提取https://www.theseason.org/nt.htm右侧圣经书籍索引及章节文本
一、优化Python爬虫(解决链接缺失+抓取章节内容)
页面右侧的圣经书籍索引属于静态结构,之前缺失链接大概率是选择器定位不准确。以下是可落地的实现步骤:
精准定位右侧索引区域
打开目标页面的浏览器开发者工具,找到右侧索引所在的HTML容器(比如特定类名的div或table),避免误抓左侧视频/论坛链接。例如,若右侧索引在class="bible-index"的表格内,就用这个选择器过滤。完整提取书籍链接
处理相对路径,用urljoin拼接成绝对URL,同时通过文本过滤排除底部无关条目(如“About”“Contact”等)。递归抓取章节文本
进入每个书籍页面,提取章节链接,再进入章节页面抓取核心文本区域(通常是id="content"或类似的容器),按「书籍文件夹-章节TXT」的结构保存。
示例代码:
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin import os base_url = "https://www.theseason.org/nt.htm" headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"} # 获取主页面并定位右侧索引容器(需根据实际页面结构调整选择器) response = requests.get(base_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") right_index = soup.find("table", attrs={"class": "nt-index"}) # 替换为实际容器选择器 # 提取目标书籍链接 book_links = [] for a in right_index.find_all("a"): link_text = a.get_text(strip=True) # 过滤无关条目,保留圣经书籍名称 if link_text in ["马太福音", "马可福音", "路加福音", "约翰福音", "使徒行传", "罗马书", "哥林多前书", "哥林多后书", "加拉太书", "以弗所书", "腓立比书", "歌罗西书", "帖撒罗尼迦前书", "帖撒罗尼迦后书", "提摩太前书", "提摩太后书", "提多书", "腓利门书", "希伯来书", "雅各书", "彼得前书", "彼得后书", "约翰一书", "约翰二书", "约翰三书", "犹大书", "启示录"]: book_url = urljoin(base_url, a["href"]) book_links.append((link_text, book_url)) # 遍历书籍,抓取章节内容 for book_name, book_url in book_links: book_dir = os.path.join("圣经文本", book_name) os.makedirs(book_dir, exist_ok=True) book_response = requests.get(book_url, headers=headers) book_soup = BeautifulSoup(book_response.text, "html.parser") # 提取章节链接(需根据书籍页面结构调整选择器) chapter_links = book_soup.find("div", id="chapter-list").find_all("a") # 替换为实际章节列表容器 for chapter in chapter_links: chapter_text = chapter.get_text(strip=True) chapter_url = urljoin(book_url, chapter["href"]) chapter_response = requests.get(chapter_url, headers=headers) chapter_soup = BeautifulSoup(chapter_response.text, "html.parser") # 提取章节核心文本 content_div = chapter_soup.find("div", id="main-content") # 替换为实际文本容器 if content_div: chapter_file = os.path.join(book_dir, f"{chapter_text}.txt") with open(chapter_file, "w", encoding="utf-8") as f: f.write(content_div.get_text(strip=True, separator="\n"))
二、WinHTTrack爬取后构建层级结构
WinHTTrack爬取的文件结构和网站一致,可通过以下方式整理:
筛选目标文件
打开爬取生成的index.html,找到右侧书籍对应的本地HTML文件路径,将这些文件单独移至临时文件夹。批量整理脚本
用Python脚本遍历本地HTML文件,按书籍-章节的层级提取文本并保存:
import os from bs4 import BeautifulSoup crawl_root = "path/to/your/winhttrack/folder" output_root = "圣经文本" os.makedirs(output_root, exist_ok=True) for root, dirs, files in os.walk(crawl_root): for file in files: if not file.endswith(".html"): continue file_path = os.path.join(root, file) with open(file_path, "r", encoding="utf-8") as f: soup = BeautifulSoup(f.read(), "html.parser") page_title = soup.title.get_text(strip=True) if soup.title else "" # 判断是否为书籍页面 if "圣经" in page_title and "第" not in page_title: book_name = page_title.split("|")[0].strip() book_dir = os.path.join(output_root, book_name) os.makedirs(book_dir, exist_ok=True) # 提取章节对应的本地文件 chapter_links = soup.find_all("a", href=True) for chapter in chapter_links: chapter_text = chapter.get_text(strip=True) if "第" not in chapter_text: continue chapter_local_path = os.path.join(root, chapter["href"]) if not os.path.exists(chapter_local_path): continue with open(chapter_local_path, "r", encoding="utf-8") as cf: chapter_soup = BeautifulSoup(cf.read(), "html.parser") content_div = chapter_soup.find("div", id="main-content") if content_div: chapter_file = os.path.join(book_dir, f"{chapter_text}.txt") with open(chapter_file, "w", encoding="utf-8") as f: f.write(content_div.get_text(strip=True, separator="\n"))
- 手动整理(小批量场景)
若书籍数量少,直接为每本书创建单独文件夹,打开章节HTML后复制核心文本,保存为对应章节名称的TXT文件即可。
三、手动高效操作技巧
- 用浏览器开发者工具右键复制右侧索引的HTML元素,粘贴到Excel中整理出所有书籍链接。
- 安装「SingleFile」浏览器插件,批量将章节页面保存为HTML,再用本地工具(如Notepad++批量转换)转成TXT。
- 将所有章节链接添加到浏览器书签,用书签批量导出工具导出后,配合批量页面保存工具抓取内容。
内容的提问来源于stack exchange,提问作者PathFinder
相关产品推荐
相关产品推荐

