You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup提取标签内文本?爬取遇阻求助

解决方法与思路

1. 定位真实内容页(核心问题)

你提供的链接是框架页(frames.php),目标文本实际嵌套在页面的<iframe>中,直接请求该链接只能拿到框架结构,无法获取里面的内容。需要先提取iframe的真实链接,再请求对应页面:

from bs4 import BeautifulSoup
import requests

# 请求框架页获取iframe地址
link = 'http://agence-prd.ansm.sante.fr/php/ecodex/frames.php?specid=61350428&typedoc=R&ref=R0390743.htm'
req = requests.get(link)
# 指定解析器,避免兼容性问题
soup = BeautifulSoup(req.content, 'html.parser')

# 提取iframe的src属性,拼接完整URL
iframe_src = soup.find('iframe')['src']
full_content_url = f'http://agence-prd.ansm.sante.fr{iframe_src}'

# 请求真实内容页
content_req = requests.get(full_content_url)
content_soup = BeautifulSoup(content_req.content, 'html.parser')

2. 提取动态class的<p>标签文本

由于<p>的class是动态变化的,无法通过class定位,可直接提取所有<p>标签,再过滤掉空文本或无效内容:

# 提取所有p标签
all_p = content_soup.find_all('p')

# 遍历过滤并输出有效文本
for p in all_p:
    # strip=True去除文本前后的空白字符
    text = p.get_text(strip=True)
    # 只输出非空的有效文本
    if text:
        print(text)

进阶优化:通过父元素缩小范围

如果页面中有很多无关的<p>标签,可以先定位包含目标文本的父容器(比如某个固定class的<div>),再在容器内提取<p>标签,减少无效内容:

# 示例:假设目标文本在class为"article-content"的div下(需根据实际网页结构调整)
target_container = content_soup.find('div', class_='article-content')
if target_container:
    target_p = target_container.find_all('p')
    for p in target_p:
        text = p.get_text(strip=True)
        if text:
            print(text)

内容的提问来源于stack exchange,提问作者yuttokb2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 16:30:49