You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页索引爬取及层级文件夹构建技术咨询

解决方案:提取https://www.theseason.org/nt.htm右侧圣经书籍索引及章节文本

一、优化Python爬虫(解决链接缺失+抓取章节内容)

页面右侧的圣经书籍索引属于静态结构,之前缺失链接大概率是选择器定位不准确。以下是可落地的实现步骤:

  1. 精准定位右侧索引区域
    打开目标页面的浏览器开发者工具,找到右侧索引所在的HTML容器(比如特定类名的div或table),避免误抓左侧视频/论坛链接。例如,若右侧索引在class="bible-index"的表格内,就用这个选择器过滤。

  2. 完整提取书籍链接
    处理相对路径,用urljoin拼接成绝对URL,同时通过文本过滤排除底部无关条目(如“About”“Contact”等)。

  3. 递归抓取章节文本
    进入每个书籍页面,提取章节链接,再进入章节页面抓取核心文本区域(通常是id="content"或类似的容器),按「书籍文件夹-章节TXT」的结构保存。

示例代码:

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
import os

base_url = "https://www.theseason.org/nt.htm"
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"}

# 获取主页面并定位右侧索引容器(需根据实际页面结构调整选择器)
response = requests.get(base_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")
right_index = soup.find("table", attrs={"class": "nt-index"})  # 替换为实际容器选择器

# 提取目标书籍链接
book_links = []
for a in right_index.find_all("a"):
    link_text = a.get_text(strip=True)
    # 过滤无关条目,保留圣经书籍名称
    if link_text in ["马太福音", "马可福音", "路加福音", "约翰福音", "使徒行传", "罗马书", "哥林多前书", "哥林多后书", "加拉太书", "以弗所书", "腓立比书", "歌罗西书", "帖撒罗尼迦前书", "帖撒罗尼迦后书", "提摩太前书", "提摩太后书", "提多书", "腓利门书", "希伯来书", "雅各书", "彼得前书", "彼得后书", "约翰一书", "约翰二书", "约翰三书", "犹大书", "启示录"]:
        book_url = urljoin(base_url, a["href"])
        book_links.append((link_text, book_url))

# 遍历书籍,抓取章节内容
for book_name, book_url in book_links:
    book_dir = os.path.join("圣经文本", book_name)
    os.makedirs(book_dir, exist_ok=True)
    
    book_response = requests.get(book_url, headers=headers)
    book_soup = BeautifulSoup(book_response.text, "html.parser")
    
    # 提取章节链接(需根据书籍页面结构调整选择器)
    chapter_links = book_soup.find("div", id="chapter-list").find_all("a")  # 替换为实际章节列表容器
    for chapter in chapter_links:
        chapter_text = chapter.get_text(strip=True)
        chapter_url = urljoin(book_url, chapter["href"])
        
        chapter_response = requests.get(chapter_url, headers=headers)
        chapter_soup = BeautifulSoup(chapter_response.text, "html.parser")
        
        # 提取章节核心文本
        content_div = chapter_soup.find("div", id="main-content")  # 替换为实际文本容器
        if content_div:
            chapter_file = os.path.join(book_dir, f"{chapter_text}.txt")
            with open(chapter_file, "w", encoding="utf-8") as f:
                f.write(content_div.get_text(strip=True, separator="\n"))

二、WinHTTrack爬取后构建层级结构

WinHTTrack爬取的文件结构和网站一致,可通过以下方式整理:

  1. 筛选目标文件
    打开爬取生成的index.html,找到右侧书籍对应的本地HTML文件路径,将这些文件单独移至临时文件夹。

  2. 批量整理脚本
    用Python脚本遍历本地HTML文件,按书籍-章节的层级提取文本并保存:

import os
from bs4 import BeautifulSoup

crawl_root = "path/to/your/winhttrack/folder"
output_root = "圣经文本"
os.makedirs(output_root, exist_ok=True)

for root, dirs, files in os.walk(crawl_root):
    for file in files:
        if not file.endswith(".html"):
            continue
        file_path = os.path.join(root, file)
        with open(file_path, "r", encoding="utf-8") as f:
            soup = BeautifulSoup(f.read(), "html.parser")
        
        page_title = soup.title.get_text(strip=True) if soup.title else ""
        # 判断是否为书籍页面
        if "圣经" in page_title and "第" not in page_title:
            book_name = page_title.split("|")[0].strip()
            book_dir = os.path.join(output_root, book_name)
            os.makedirs(book_dir, exist_ok=True)
            
            # 提取章节对应的本地文件
            chapter_links = soup.find_all("a", href=True)
            for chapter in chapter_links:
                chapter_text = chapter.get_text(strip=True)
                if "第" not in chapter_text:
                    continue
                chapter_local_path = os.path.join(root, chapter["href"])
                if not os.path.exists(chapter_local_path):
                    continue
                
                with open(chapter_local_path, "r", encoding="utf-8") as cf:
                    chapter_soup = BeautifulSoup(cf.read(), "html.parser")
                    content_div = chapter_soup.find("div", id="main-content")
                    if content_div:
                        chapter_file = os.path.join(book_dir, f"{chapter_text}.txt")
                        with open(chapter_file, "w", encoding="utf-8") as f:
                            f.write(content_div.get_text(strip=True, separator="\n"))
  1. 手动整理(小批量场景)
    若书籍数量少,直接为每本书创建单独文件夹,打开章节HTML后复制核心文本,保存为对应章节名称的TXT文件即可。

三、手动高效操作技巧

  • 用浏览器开发者工具右键复制右侧索引的HTML元素,粘贴到Excel中整理出所有书籍链接。
  • 安装「SingleFile」浏览器插件,批量将章节页面保存为HTML,再用本地工具(如Notepad++批量转换)转成TXT。
  • 将所有章节链接添加到浏览器书签,用书签批量导出工具导出后,配合批量页面保存工具抓取内容。

内容的提问来源于stack exchange,提问作者PathFinder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 22:14:56