网页抓取场景中如何精准提取指定目标文本?
解决方案
首先,问题出在你直接抓取<br>标签的文本——<br>是换行标签,本身没有文本内容,所以输出为空。要提取目标内容,你需要获取td.testo_normale节点下的所有文本片段,再筛选出需要的内容。
方法一:利用stripped_strings提取文本片段
stripped_strings会返回节点下所有非空白的文本内容,你可以遍历这些内容并筛选目标:
from bs4 import BeautifulSoup import requests r = requests.get('https://www.deswater.com/vol.php?vol=1&oth=1|1-3|January|2009') soup = BeautifulSoup(r.content, 'html.parser') target_td = soup.find('td', class_='testo_normale') # 获取所有非空白文本片段 text_fragments = list(target_td.stripped_strings) # 筛选目标内容 for fragment in text_fragments: if fragment == 'Water engineering' or '(2009)' in fragment: print(fragment)
方法二:分割文本内容
先获取节点的全部文本,再按换行符分割后提取:
from bs4 import BeautifulSoup import requests r = requests.get('https://www.deswater.com/vol.php?vol=1&oth=1|1-3|January|2009') soup = BeautifulSoup(r.content, 'html.parser') target_td = soup.find('td', class_='testo_normale') # 获取文本并按换行分割,过滤空白行 all_lines = [line.strip() for line in target_td.get_text('\n').split('\n') if line.strip()] # 提取目标内容 water_eng = all_lines[1] # 对应"Water engineering" issue_info = all_lines[2] # 对应"1(2009)1-6" print(water_eng) print(issue_info)
说明
两种方法都能精准提取你需要的内容:
- 方法一更灵活,不用依赖内容的固定位置;
- 方法二适合你确定目标内容在固定行的场景。
内容的提问来源于stack exchange,提问作者user17356493
相关产品推荐
相关产品推荐

