如何使用BeautifulSoup提取id前缀为statement的不同div标签内的文本
解决方法
BeautifulSoup 的属性筛选参数支持传入正则表达式做模糊匹配,你可以直接用正则匹配id以statement开头的div标签,同时增加非空判断避免空对象报错,修改后的代码如下:
import requests from bs4 import BeautifulSoup import re limit = 100 url = f'https://www.counselingcalifornia.com/cc/cgi-bin/utilities.dll/customlist?FIRSTNAME=~&LASTNAME=~&ZIP=&DONORCLASSSTT=&_MULTIPLE_INSURANCE=&HASPHOTOFLG=&_MULTIPLE_EMPHASIS=ÐNIC=&_MULTIPLE_LANGUAGE=ENG&QNAME=THERAPISTLIST&WMT=NONE&WNR=NONE&WHP=therapistHeader.htm&WBP=therapistList.htm&RANGE=1%2F{limit}&SORT=LASTNAME' headers = {'User-Agent': 'Mozilla/5.0 (Linux; Android 6.0; Nexus 5 Build/MRA58N) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Mobile Safari/537.36'} response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') rows = soup.find_all('div', {'class':'row'}) for row in rows: # 用正则匹配id以statement开头的div des = row.find('div', id=re.compile(r'^statement')) if des: # 非空判断,避免不存在对应标签时报错 print(des.text.strip())
改动说明
- 将原代码中固定的
id='statement80863'替换为id=re.compile(r'^statement'),正则中的^符号代表匹配字符串开头,完美适配所有前缀为statement的动态生成id - 新增非空判断逻辑,避免部分行下没有对应div时,调用
.text属性触发AttributeError报错
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

