You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解析大文件时,如何在遍历div中使用.find()获取<a>标签内容?

嘿,我看你在爬PMU的赛事页面时遇到麻烦了,代码跑不起来对吧?先瞅一眼你给出的代码片段——class_...这里明显没写完,这肯定会直接报错!另外还有个关键坑:这类体育博彩网站的赛事内容几乎都是JavaScript动态渲染的,用requests直接GET只能拿到空壳静态HTML,根本抓不到那些动态加载出来的目标div和a标签内容。

下面给你一步步解决的方案:

1. 先补全基础代码(静态页面场景)

如果页面是纯静态的(虽然PMU大概率不是,但先给你正确的写法参考),你需要明确目标div的class名称(可以用浏览器F12开发者工具查看),然后遍历每个div下的a标签:

from bs4 import BeautifulSoup
import requests

page_url = 'https://paris-sportifs.pmu.fr/'
# 加请求头模拟浏览器,避免被反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
page = requests.get(page_url, headers=headers)
soup = BeautifulSoup(page.text, 'html.parser')

# 替换成你实际要找的div的class!如果class有空格,写成列表形式比如['card', 'event-item']
with open('pmu.html', 'a+', encoding='utf-8') as file:
    for div in soup.find_all('div', class_='your-target-div-class'):
        # 遍历当前div下所有的a标签
        for a_tag in div.find_all('a'):
            # 获取a标签的所有文本(自动合并子元素文本)
            a_text = a_tag.get_text(strip=True, separator=' ')
            # 获取a标签的链接(没有的话返回空字符串)
            a_link = a_tag.get('href', '')
            # 获取a标签的完整HTML内容(包括子元素)
            a_html = str(a_tag)
            # 写入文件,格式可按需调整
            file.write(f"链接: {a_link}\n文本内容: {a_text}\n完整HTML:\n{a_html}\n---分割线---\n")
2. 处理动态加载的核心问题(PMU的真实情况)

PMU的赛事数据都是JS动态加载的,requests拿不到,这时候得用能执行JS的工具,给你两个常用方案:

方案A:用Selenium(最稳定)

需要先安装Selenium和对应浏览器的驱动(比如ChromeDriver):

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time

# 配置Chrome无头模式(后台运行不弹浏览器窗口)
chrome_options = Options()
chrome_options.add_argument('--headless=new')
chrome_options.add_argument('--disable-gpu')
chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36')

driver = webdriver.Chrome(options=chrome_options)
driver.get('https://paris-sportifs.pmu.fr/')
# 等待页面加载完成,根据网络情况调整等待时间(比如3-5秒)
time.sleep(3)

# 获取JS渲染后的完整页面源码
page_source = driver.page_source
soup = BeautifulSoup(page_source, 'html.parser')

# 后续逻辑和静态页面一致,替换class即可
with open('pmu.html', 'a+', encoding='utf-8') as file:
    for div in soup.find_all('div', class_='your-target-div-class'):
        for a_tag in div.find_all('a'):
            a_text = a_tag.get_text(strip=True, separator=' ')
            a_link = a_tag.get('href', '')
            a_html = str(a_tag)
            file.write(f"链接: {a_link}\n文本内容: {a_text}\n完整HTML:\n{a_html}\n---分割线---\n")

# 关闭浏览器驱动
driver.quit()

方案B:用requests-html(更轻量)

不需要装浏览器驱动,直接用库自带的JS渲染:

from bs4 import BeautifulSoup
from requests_html import HTMLSession

session = HTMLSession()
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
r = session.get('https://paris-sportifs.pmu.fr/', headers=headers)
# 执行JS渲染页面,sleep参数是等待加载的时间
r.html.render(sleep=3)

soup = BeautifulSoup(r.html.html, 'html.parser')

# 同样替换目标div的class即可
with open('pmu.html', 'a+', encoding='utf-8') as file:
    for div in soup.find_all('div', class_='your-target-div-class'):
        for a_tag in div.find_all('a'):
            a_text = a_tag.get_text(strip=True, separator=' ')
            a_link = a_tag.get('href', '')
            a_html = str(a_tag)
            file.write(f"链接: {a_link}\n文本内容: {a_text}\n完整HTML:\n{a_html}\n---分割线---\n")
3. 几个重要提醒
  • 替换class名称:一定要把代码里的your-target-div-class换成你实际要抓取的div的class,用浏览器F12选元素就能看到。
  • 反爬注意:PMU有反爬机制,不要频繁请求,最好加随机延迟(比如用time.sleep(random.uniform(1,3))),避免被封IP。
  • 编码问题:写入文件时一定要加encoding='utf-8',不然容易出现乱码。

内容的提问来源于stack exchange,提问作者Elsha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:06:50