如何用Python+requests+bs4获取指定class的div的onclick属性?
解决方案:遍历onclick属性并爬取对应页面内容
方案1:基础提取与串行遍历
这是对你初步想法的细化实现,核心是提取onclick中的URL并逐个请求:
import requests from bs4 import BeautifulSoup import re # 配置基础参数 target_url = "你的目标网页URL" headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} # 获取目标页面并解析 response = requests.get(target_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 定位所有目标class的div元素 target_divs = soup.find_all("div", class_="你的目标class名称") # 用正则匹配onclick中的URL(适配单/双引号包裹的格式) url_regex = re.compile(r"['\"](https?://.*?|/.*?)['\"]") # 遍历处理每个div for div in target_divs: onclick_value = div.get("onclick") if not onclick_value: continue # 跳过无onclick属性的div # 提取URL url_match = url_regex.search(onclick_value) if url_match: raw_url = url_match.group(1) # 处理相对URL,拼接成绝对地址 detail_url = requests.compat.urljoin(target_url, raw_url) # 请求详情页并处理数据 try: detail_resp = requests.get(detail_url, headers=headers) # 这里写你的数据解析/保存逻辑,比如: # detail_soup = BeautifulSoup(detail_resp.text, "html.parser") # save_data(detail_soup) print(f"已处理详情页:{detail_url}") except Exception as e: print(f"处理{detail_url}失败:{str(e)}")
方案2:处理复杂JS函数式onclick
如果onclick不是直接的跳转语句,而是调用JS函数(比如goDetail(12345)),可以解析函数参数拼接URL:
# 承接上面的target_divs遍历逻辑 func_regex = re.compile(r"goDetail\((.*?)\)") # 替换为实际的函数名 for div in target_divs: onclick_value = div.get("onclick") if not onclick_value: continue func_match = func_regex.search(onclick_value) if func_match: # 提取并清理参数 params = [p.strip().strip("'\"") for p in func_match.group(1).split(",")] # 假设第一个参数是ID,拼接详情页URL detail_url = f"http://example.com/detail?id={params[0]}" # 替换为实际URL规则 # 后续请求逻辑同方案1
方案3:并发请求提升效率
面对500个页面,串行请求效率太低,用线程池并发请求可以大幅提速(注意控制并发数,避免触发反爬):
import requests from bs4 import BeautifulSoup import re from concurrent.futures import ThreadPoolExecutor # 先收集所有详情页URL target_url = "你的目标网页URL" headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} response = requests.get(target_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") target_divs = soup.find_all("div", class_="你的目标class名称") url_regex = re.compile(r"['\"](https?://.*?|/.*?)['\"]") detail_urls = [] for div in target_divs: onclick_value = div.get("onclick") if not onclick_value: continue url_match = url_regex.search(onclick_value) if url_match: raw_url = url_match.group(1) detail_url = requests.compat.urljoin(target_url, raw_url) detail_urls.append(detail_url) # 定义单页爬取函数 def crawl_single_page(url): try: resp = requests.get(url, headers=headers) # 数据处理逻辑 # detail_soup = BeautifulSoup(resp.text, "html.parser") return (url, True) except Exception as e: print(f"失败:{url} - {str(e)}") return (url, False) # 启动线程池并发请求(max_workers建议10-20,根据网站反爬强度调整) with ThreadPoolExecutor(max_workers=15) as executor: results = executor.map(crawl_single_page, detail_urls) # 可选:统计结果 success_count = sum(1 for _, success in results if success) print(f"爬取完成,成功{success_count}个,失败{len(detail_urls)-success_count}个")
关键注意事项
- 必须添加
User-Agent等请求头,模拟浏览器行为,避免被网站直接拦截。 - 相对URL一定要用
requests.compat.urljoin拼接,否则会请求错误的地址。 - 可以加入随机延迟(
import time; import random; time.sleep(random.uniform(0.5, 2))),降低被反爬的概率。 - 异常处理不能少,避免单个页面请求失败导致整个程序崩溃。
内容的提问来源于stack exchange,提问作者SnakePotato
相关产品推荐
相关产品推荐

