Python BeautifulSoup提取指定caption对应href及批量URL处理问题
批量提取符合条件的页面链接解决方案
核心逻辑
要实现需求,关键是通过HTML结构层级定位元素:
- 先找到所有带
caption类的div元素 - 检查其文本内容是否包含"English"
- 若符合条件,找到它的直接父元素(带
cover类的a标签) - 提取该a标签的
href属性,补全为完整URL后保存
所需依赖
先安装两个Python库,打开终端运行:
pip install requests beautifulsoup4
完整代码实现
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin def extract_target_links(url): # 模拟浏览器请求头,避免被网站拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } try: # 发送请求获取页面内容 response = requests.get(url, headers=headers) response.raise_for_status() # 捕获请求失败的异常 soup = BeautifulSoup(response.text, "html.parser") target_links = [] # 遍历所有caption元素 for caption in soup.find_all("div", class_="caption"): # 清理文本后检查是否包含"English" caption_text = caption.get_text(strip=True) if "English" in caption_text: # 找到父级的cover类a标签 cover_link = caption.find_parent("a", class_="cover") if cover_link and "href" in cover_link.attrs: raw_href = cover_link["href"] # 补全相对路径为完整URL full_link = urljoin(url, raw_href) target_links.append(full_link) return target_links except Exception as e: print(f"处理URL {url} 时出错:{str(e)}") return [] def batch_process(txt_file_path, output_file="extracted_links.txt"): # 读取txt文件中的所有URL(每行一个) with open(txt_file_path, "r", encoding="utf-8") as f: urls = [line.strip() for line in f if line.strip()] # 跳过空行 all_links = [] # 逐个处理每个URL for idx, url in enumerate(urls, 1): print(f"正在处理第 {idx}/{len(urls)} 个URL:{url}") links = extract_target_links(url) all_links.extend(links) # 将结果写入输出文件 with open(output_file, "w", encoding="utf-8") as f: for link in all_links: f.write(f"{link}\n") print(f"处理完成!共提取到 {len(all_links)} 条符合条件的链接,已保存到 {output_file}") # ------------------- 使用示例 ------------------- # 1. 处理单个URL # single_url = "https://your-target-site.com" # links = extract_target_links(single_url) # print("提取的链接:", links) # 2. 批量处理txt文件中的URL(将所有URL放在urls.txt中,每行一个) batch_process("urls.txt")
代码关键点说明
- 请求头模拟:通过
headers设置浏览器标识,避免被目标网站的反爬机制拦截 - 元素定位:用
find_parent精准定位caption的父级a标签,比遍历上级元素更可靠 - URL补全:用
urljoin自动将相对路径(如/g/987654/)转换为完整URL,确保链接可直接访问 - 异常处理:捕获请求失败、元素不存在等异常,保证批量处理时不会因单个URL出错而中断
- 空行过滤:读取txt文件时自动跳过空行,只处理有效URL
内容的提问来源于stack exchange,提问作者Kirizu
相关产品推荐
相关产品推荐

