Python未知网站结构下的网页爬取方案咨询(新手向)
批量搜索指定文档的穷举式方案(针对200+站点)
针对你需要处理大量站点、避免逐个分析结构的需求,以下是几种实用的穷举式搜索思路,结合Python实现要点:
一、基于路径关键词的定向穷举
利用你观察到的路径相似性,提取目标文档的核心关键词,生成可能的URL组合去验证,适合路径有规律的站点(比如你的两个示例)。
核心步骤:
- 提炼关键词:从已知示例中提取目标文档的核心标识,比如:
- 文档名变体:
carta-anual、governanca-corporativa、consad - 站点通用路径:
transparencia、acesso-a-informacao - 附加信息:年份(2019-2023)、文件后缀(
.pdf、.docx)
- 文档名变体:
- 生成路径组合:将关键词、年份、后缀按常见站点路径格式组合,生成待测试URL
- 批量验证下载:用
requests请求URL,通过状态码和响应类型判断是否为目标文档
代码示例:
import requests import time # 配置参数 base_urls = ["https://ceagesp.gov.br/", "https://www.amazul.mar.mil.br/"] core_keywords = ["carta-anual", "governanca", "transparencia", "consad"] years = ["2019", "2020", "2021", "2022", "2023"] extensions = [".pdf", "", "/"] for base in base_urls: base_domain = base.split("//")[1].rstrip("/").replace(".", "_") # 生成多种路径格式 for kw1 in core_keywords: for kw2 in core_keywords: if kw1 == kw2: continue for year in years: for ext in extensions: test_urls = [ f"{base.rstrip('/')}/{kw1}/{kw2}/{year}{ext}", f"{base.rstrip('/')}/{kw1}-{kw2}-{year}{ext}", f"{base.rstrip('/')}/transparencia/{kw1}-{year}{ext}" ] for url in test_urls: try: # 先用HEAD请求减少流量 resp = requests.head(url, timeout=5, headers={"User-Agent": "Mozilla/5.0"}) if resp.status_code == 200: if "application/pdf" in resp.headers.get("Content-Type", ""): print(f"找到目标文档: {url}") # 下载文件 file_name = f"{base_domain}_{year}_{kw1}.pdf" with open(file_name, "wb") as f: f.write(requests.get(url, headers={"User-Agent": "Mozilla/5.0"}).content) time.sleep(1) # 避免请求过于频繁 except requests.exceptions.RequestException: continue
二、站内链接广度遍历(半穷举)
如果路径规律不明显,可以爬取站点内的所有可访问链接,筛选包含目标关键词的内容,适合结构复杂但关键词明确的站点。
核心步骤:
- 爬取站内链接:用
BeautifulSoup解析页面,提取所有站内<a>标签链接 - 关键词筛选:检查链接URL或页面标题是否包含目标文档关键词
- 文档验证:对筛选出的链接,验证是否为目标格式的文档
代码示例:
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin, urlparse from concurrent.futures import ThreadPoolExecutor def check_and_download(url, target_keywords): try: resp = requests.head(url, timeout=5, headers={"User-Agent": "Mozilla/5.0"}) if resp.status_code == 200 and "application/pdf" in resp.headers.get("Content-Type", ""): if any(kw in url.lower() for kw in target_keywords): print(f"找到目标文档: {url}") file_name = url.split("/")[-1] if url.split("/")[-1] else f"{urlparse(url).netloc}_document.pdf" with open(file_name, "wb") as f: f.write(requests.get(url, headers={"User-Agent": "Mozilla/5.0"}).content) except requests.exceptions.RequestException: pass def crawl_site(base_url, target_keywords, max_depth=3): visited = set() queue = [(base_url, 0)] while queue: current_url, depth = queue.pop(0) if depth > max_depth or current_url in visited: continue visited.add(current_url) try: resp = requests.get(current_url, timeout=5, headers={"User-Agent": "Mozilla/5.0"}) if resp.status_code != 200: continue soup = BeautifulSoup(resp.text, "html.parser") links = [] for a_tag in soup.find_all("a", href=True): full_url = urljoin(base_url, a_tag["href"]) if urlparse(full_url).netloc == urlparse(base_url).netloc and full_url not in visited: links.append(full_url) if depth < max_depth: queue.append((full_url, depth + 1)) # 多线程验证链接 with ThreadPoolExecutor(max_workers=5) as executor: executor.map(lambda url: check_and_download(url, target_keywords), links) except requests.exceptions.RequestException: continue # 使用示例 target_keywords = ["carta anual", "governanca corporativa", "consad"] for base in ["https://ceagesp.gov.br/", "https://www.amazul.mar.mil.br/"]: crawl_site(base, target_keywords)
三、关键优化点
- 反爬规避:设置浏览器
User-Agent,添加请求间隔,必要时使用代理IP - 效率提升:用
HEAD请求替代GET做初步验证,多线程/多进程处理批量站点 - 关键词扩展:收集目标文档的所有可能变体(比如缩写、不同语言拼写),提升命中率
- 结果校验:下载后可检查文件大小、文件名或内容关键词,避免误下载
内容的提问来源于stack exchange,提问作者Larissa Kelmer
相关产品推荐
相关产品推荐

