You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取场景中如何精准提取指定目标文本?

解决方案

首先,问题出在你直接抓取<br>标签的文本——<br>是换行标签,本身没有文本内容,所以输出为空。要提取目标内容,你需要获取td.testo_normale节点下的所有文本片段,再筛选出需要的内容。

方法一:利用stripped_strings提取文本片段

stripped_strings会返回节点下所有非空白的文本内容,你可以遍历这些内容并筛选目标:

from bs4 import BeautifulSoup
import requests

r = requests.get('https://www.deswater.com/vol.php?vol=1&oth=1|1-3|January|2009')
soup = BeautifulSoup(r.content, 'html.parser')
target_td = soup.find('td', class_='testo_normale')

# 获取所有非空白文本片段
text_fragments = list(target_td.stripped_strings)

# 筛选目标内容
for fragment in text_fragments:
    if fragment == 'Water engineering' or '(2009)' in fragment:
        print(fragment)

方法二:分割文本内容

先获取节点的全部文本,再按换行符分割后提取:

from bs4 import BeautifulSoup
import requests

r = requests.get('https://www.deswater.com/vol.php?vol=1&oth=1|1-3|January|2009')
soup = BeautifulSoup(r.content, 'html.parser')
target_td = soup.find('td', class_='testo_normale')

# 获取文本并按换行分割,过滤空白行
all_lines = [line.strip() for line in target_td.get_text('\n').split('\n') if line.strip()]

# 提取目标内容
water_eng = all_lines[1]  # 对应"Water engineering"
issue_info = all_lines[2] # 对应"1(2009)1-6"
print(water_eng)
print(issue_info)

说明

两种方法都能精准提取你需要的内容:

  • 方法一更灵活,不用依赖内容的固定位置;
  • 方法二适合你确定目标内容在固定行的场景。

内容的提问来源于stack exchange,提问作者user17356493

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 18:15:51