You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python未知网站结构下的网页爬取方案咨询(新手向)

批量搜索指定文档的穷举式方案(针对200+站点)

针对你需要处理大量站点、避免逐个分析结构的需求,以下是几种实用的穷举式搜索思路,结合Python实现要点:

一、基于路径关键词的定向穷举

利用你观察到的路径相似性,提取目标文档的核心关键词,生成可能的URL组合去验证,适合路径有规律的站点(比如你的两个示例)。

核心步骤:

  1. 提炼关键词:从已知示例中提取目标文档的核心标识,比如:
    • 文档名变体:carta-anual、governanca-corporativa、consad
    • 站点通用路径:transparencia、acesso-a-informacao
    • 附加信息:年份(2019-2023)、文件后缀(.pdf、.docx)
  2. 生成路径组合:将关键词、年份、后缀按常见站点路径格式组合,生成待测试URL
  3. 批量验证下载:用requests请求URL,通过状态码和响应类型判断是否为目标文档

代码示例:

import requests
import time

# 配置参数
base_urls = ["https://ceagesp.gov.br/", "https://www.amazul.mar.mil.br/"]
core_keywords = ["carta-anual", "governanca", "transparencia", "consad"]
years = ["2019", "2020", "2021", "2022", "2023"]
extensions = [".pdf", "", "/"]

for base in base_urls:
    base_domain = base.split("//")[1].rstrip("/").replace(".", "_")
    # 生成多种路径格式
    for kw1 in core_keywords:
        for kw2 in core_keywords:
            if kw1 == kw2:
                continue
            for year in years:
                for ext in extensions:
                    test_urls = [
                        f"{base.rstrip('/')}/{kw1}/{kw2}/{year}{ext}",
                        f"{base.rstrip('/')}/{kw1}-{kw2}-{year}{ext}",
                        f"{base.rstrip('/')}/transparencia/{kw1}-{year}{ext}"
                    ]
                    for url in test_urls:
                        try:
                            # 先用HEAD请求减少流量
                            resp = requests.head(url, timeout=5, headers={"User-Agent": "Mozilla/5.0"})
                            if resp.status_code == 200:
                                if "application/pdf" in resp.headers.get("Content-Type", ""):
                                    print(f"找到目标文档: {url}")
                                    # 下载文件
                                    file_name = f"{base_domain}_{year}_{kw1}.pdf"
                                    with open(file_name, "wb") as f:
                                        f.write(requests.get(url, headers={"User-Agent": "Mozilla/5.0"}).content)
                                time.sleep(1)  # 避免请求过于频繁
                        except requests.exceptions.RequestException:
                            continue

二、站内链接广度遍历(半穷举)

如果路径规律不明显,可以爬取站点内的所有可访问链接,筛选包含目标关键词的内容,适合结构复杂但关键词明确的站点。

核心步骤:

  1. 爬取站内链接:用BeautifulSoup解析页面,提取所有站内<a>标签链接
  2. 关键词筛选:检查链接URL或页面标题是否包含目标文档关键词
  3. 文档验证:对筛选出的链接,验证是否为目标格式的文档

代码示例:

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse
from concurrent.futures import ThreadPoolExecutor

def check_and_download(url, target_keywords):
    try:
        resp = requests.head(url, timeout=5, headers={"User-Agent": "Mozilla/5.0"})
        if resp.status_code == 200 and "application/pdf" in resp.headers.get("Content-Type", ""):
            if any(kw in url.lower() for kw in target_keywords):
                print(f"找到目标文档: {url}")
                file_name = url.split("/")[-1] if url.split("/")[-1] else f"{urlparse(url).netloc}_document.pdf"
                with open(file_name, "wb") as f:
                    f.write(requests.get(url, headers={"User-Agent": "Mozilla/5.0"}).content)
    except requests.exceptions.RequestException:
        pass

def crawl_site(base_url, target_keywords, max_depth=3):
    visited = set()
    queue = [(base_url, 0)]
    
    while queue:
        current_url, depth = queue.pop(0)
        if depth > max_depth or current_url in visited:
            continue
        visited.add(current_url)
        
        try:
            resp = requests.get(current_url, timeout=5, headers={"User-Agent": "Mozilla/5.0"})
            if resp.status_code != 200:
                continue
            
            soup = BeautifulSoup(resp.text, "html.parser")
            links = []
            for a_tag in soup.find_all("a", href=True):
                full_url = urljoin(base_url, a_tag["href"])
                if urlparse(full_url).netloc == urlparse(base_url).netloc and full_url not in visited:
                    links.append(full_url)
                    if depth < max_depth:
                        queue.append((full_url, depth + 1))
            
            # 多线程验证链接
            with ThreadPoolExecutor(max_workers=5) as executor:
                executor.map(lambda url: check_and_download(url, target_keywords), links)
        except requests.exceptions.RequestException:
            continue

# 使用示例
target_keywords = ["carta anual", "governanca corporativa", "consad"]
for base in ["https://ceagesp.gov.br/", "https://www.amazul.mar.mil.br/"]:
    crawl_site(base, target_keywords)

三、关键优化点

  • 反爬规避:设置浏览器User-Agent,添加请求间隔,必要时使用代理IP
  • 效率提升:用HEAD请求替代GET做初步验证,多线程/多进程处理批量站点
  • 关键词扩展:收集目标文档的所有可能变体(比如缩写、不同语言拼写),提升命中率
  • 结果校验:下载后可检查文件大小、文件名或内容关键词,避免误下载

内容的提问来源于stack exchange,提问作者Larissa Kelmer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 16:59:58