使用BeautifulSoup如何提取相同class的article标签下第二个p标签文本
解决思路
你当前的问题根源是直接从整个文档根节点查找article > p,会命中页面所有article标签下的p元素,索引取值的方式不仅不稳定还会拿到无关内容。你已经通过class属性定位到了摘要所属的唯一article节点,直接在这个节点范围内查找下属的p标签即可。
修正后代码
import requests from bs4 import BeautifulSoup tempURL = 'http://www.globalbigdataconference.com/santa-clara/global-artificial-intelligence-virtual-conference-125/speaker-details/aaron-burciaga-114059.html' page = requests.get(tempURL) soup = BeautifulSoup(page.content, 'lxml') Tag = soup.find_all('h4', class_ = 'clearfix Roboto-Medium font13 sbl-t t-b-m0') Value = soup.find_all('h4', class_ = 'clearfix Roboto-Medium font15 sbl-t t-b-m0 dks-t l-h20') # 定位到摘要所属的article容器 topic_article = soup.find( 'article', class_ = 'clearfix font14 dkg-t Roboto-Regular t-p15 l-h26') # 仅在该article容器内查找p标签,避免命中其他无关节点 abstract = topic_article.find('p').get_text().strip() print(abstract)
补充说明
如果该article下存在多个p标签,你可以根据内容特征进一步过滤,比如过滤掉空文本的p标签,或者按顺序取第一个非空p标签即可,不需要依赖全局索引,适配性更强。
内容的提问来源于stack exchange,提问作者Raspberry Lemon
相关产品推荐
相关产品推荐

