You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+requests+bs4获取指定class的div的onclick属性?

解决方案:遍历onclick属性并爬取对应页面内容

方案1:基础提取与串行遍历

这是对你初步想法的细化实现,核心是提取onclick中的URL并逐个请求:

import requests
from bs4 import BeautifulSoup
import re

# 配置基础参数
target_url = "你的目标网页URL"
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}

# 获取目标页面并解析
response = requests.get(target_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# 定位所有目标class的div元素
target_divs = soup.find_all("div", class_="你的目标class名称")

# 用正则匹配onclick中的URL(适配单/双引号包裹的格式)
url_regex = re.compile(r"['\"](https?://.*?|/.*?)['\"]")

# 遍历处理每个div
for div in target_divs:
    onclick_value = div.get("onclick")
    if not onclick_value:
        continue  # 跳过无onclick属性的div
    
    # 提取URL
    url_match = url_regex.search(onclick_value)
    if url_match:
        raw_url = url_match.group(1)
        # 处理相对URL,拼接成绝对地址
        detail_url = requests.compat.urljoin(target_url, raw_url)
        
        # 请求详情页并处理数据
        try:
            detail_resp = requests.get(detail_url, headers=headers)
            # 这里写你的数据解析/保存逻辑,比如:
            # detail_soup = BeautifulSoup(detail_resp.text, "html.parser")
            # save_data(detail_soup)
            print(f"已处理详情页:{detail_url}")
        except Exception as e:
            print(f"处理{detail_url}失败:{str(e)}")

方案2:处理复杂JS函数式onclick

如果onclick不是直接的跳转语句,而是调用JS函数(比如goDetail(12345)),可以解析函数参数拼接URL:

# 承接上面的target_divs遍历逻辑
func_regex = re.compile(r"goDetail\((.*?)\)")  # 替换为实际的函数名

for div in target_divs:
    onclick_value = div.get("onclick")
    if not onclick_value:
        continue
    
    func_match = func_regex.search(onclick_value)
    if func_match:
        # 提取并清理参数
        params = [p.strip().strip("'\"") for p in func_match.group(1).split(",")]
        # 假设第一个参数是ID,拼接详情页URL
        detail_url = f"http://example.com/detail?id={params[0]}"  # 替换为实际URL规则
        
        # 后续请求逻辑同方案1

方案3:并发请求提升效率

面对500个页面,串行请求效率太低,用线程池并发请求可以大幅提速(注意控制并发数,避免触发反爬):

import requests
from bs4 import BeautifulSoup
import re
from concurrent.futures import ThreadPoolExecutor

# 先收集所有详情页URL
target_url = "你的目标网页URL"
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}

response = requests.get(target_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")
target_divs = soup.find_all("div", class_="你的目标class名称")
url_regex = re.compile(r"['\"](https?://.*?|/.*?)['\"]")

detail_urls = []
for div in target_divs:
    onclick_value = div.get("onclick")
    if not onclick_value:
        continue
    url_match = url_regex.search(onclick_value)
    if url_match:
        raw_url = url_match.group(1)
        detail_url = requests.compat.urljoin(target_url, raw_url)
        detail_urls.append(detail_url)

# 定义单页爬取函数
def crawl_single_page(url):
    try:
        resp = requests.get(url, headers=headers)
        # 数据处理逻辑
        # detail_soup = BeautifulSoup(resp.text, "html.parser")
        return (url, True)
    except Exception as e:
        print(f"失败:{url} - {str(e)}")
        return (url, False)

# 启动线程池并发请求(max_workers建议10-20,根据网站反爬强度调整)
with ThreadPoolExecutor(max_workers=15) as executor:
    results = executor.map(crawl_single_page, detail_urls)

# 可选:统计结果
success_count = sum(1 for _, success in results if success)
print(f"爬取完成,成功{success_count}个,失败{len(detail_urls)-success_count}个")

关键注意事项

  • 必须添加User-Agent等请求头,模拟浏览器行为,避免被网站直接拦截。
  • 相对URL一定要用requests.compat.urljoin拼接,否则会请求错误的地址。
  • 可以加入随机延迟(import time; import random; time.sleep(random.uniform(0.5, 2))),降低被反爬的概率。
  • 异常处理不能少,避免单个页面请求失败导致整个程序崩溃。

内容的提问来源于stack exchange,提问作者SnakePotato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 10:52:22