You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup提取指定caption对应href及批量URL处理问题

批量提取符合条件的页面链接解决方案

核心逻辑

要实现需求,关键是通过HTML结构层级定位元素:

  • 先找到所有带caption类的div元素
  • 检查其文本内容是否包含"English"
  • 若符合条件,找到它的直接父元素(带cover类的a标签)
  • 提取该a标签的href属性,补全为完整URL后保存

所需依赖

先安装两个Python库,打开终端运行:

pip install requests beautifulsoup4

完整代码实现

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

def extract_target_links(url):
    # 模拟浏览器请求头,避免被网站拦截
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    try:
        # 发送请求获取页面内容
        response = requests.get(url, headers=headers)
        response.raise_for_status()  # 捕获请求失败的异常
        soup = BeautifulSoup(response.text, "html.parser")
        
        target_links = []
        # 遍历所有caption元素
        for caption in soup.find_all("div", class_="caption"):
            # 清理文本后检查是否包含"English"
            caption_text = caption.get_text(strip=True)
            if "English" in caption_text:
                # 找到父级的cover类a标签
                cover_link = caption.find_parent("a", class_="cover")
                if cover_link and "href" in cover_link.attrs:
                    raw_href = cover_link["href"]
                    # 补全相对路径为完整URL
                    full_link = urljoin(url, raw_href)
                    target_links.append(full_link)
        return target_links
    except Exception as e:
        print(f"处理URL {url} 时出错:{str(e)}")
        return []

def batch_process(txt_file_path, output_file="extracted_links.txt"):
    # 读取txt文件中的所有URL(每行一个)
    with open(txt_file_path, "r", encoding="utf-8") as f:
        urls = [line.strip() for line in f if line.strip()]  # 跳过空行
    
    all_links = []
    # 逐个处理每个URL
    for idx, url in enumerate(urls, 1):
        print(f"正在处理第 {idx}/{len(urls)} 个URL:{url}")
        links = extract_target_links(url)
        all_links.extend(links)
    
    # 将结果写入输出文件
    with open(output_file, "w", encoding="utf-8") as f:
        for link in all_links:
            f.write(f"{link}\n")
    
    print(f"处理完成!共提取到 {len(all_links)} 条符合条件的链接,已保存到 {output_file}")

# ------------------- 使用示例 -------------------
# 1. 处理单个URL
# single_url = "https://your-target-site.com"
# links = extract_target_links(single_url)
# print("提取的链接:", links)

# 2. 批量处理txt文件中的URL(将所有URL放在urls.txt中,每行一个)
batch_process("urls.txt")

代码关键点说明

  • 请求头模拟:通过headers设置浏览器标识,避免被目标网站的反爬机制拦截
  • 元素定位:用find_parent精准定位caption的父级a标签,比遍历上级元素更可靠
  • URL补全:用urljoin自动将相对路径(如/g/987654/)转换为完整URL,确保链接可直接访问
  • 异常处理:捕获请求失败、元素不存在等异常,保证批量处理时不会因单个URL出错而中断
  • 空行过滤:读取txt文件时自动跳过空行,只处理有效URL

内容的提问来源于stack exchange,提问作者Kirizu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 11:31:01