如何在BeautifulSoup中通过-soup-contains选中节点后获取后续两个相邻兄弟节点
错误原因
你当前的CSS选择器逻辑有误:包含PPP文本的th元素本身嵌套在<tr>标签内,th的同级元素是同单元格行的<td>,而非下一行的<tr>。你写的+ tr是在查找和th同级的<tr>,自然匹配不到正确结果。
正确实现方案
先定位到包含带PPP文本的th的整行tr,再用兄弟选择器匹配该行后续的所有<tr>,就能拿到总GDP、人均GDP对应的两行数据,可运行代码如下:
import requests from bs4 import BeautifulSoup site = "http://en.wikipedia.org/wiki/Brazil" country = requests.get(site) countryPage = BeautifulSoup(country.content, "html.parser") infoBox = countryPage.find("table", class_="infobox ib-country vcard") # 定位GDP PPP所在行,再选后续的2行数据 gdp_ppp_rows = infoBox.select('tr:has(th:-soup-contains("PPP")) ~ tr') # 第一行是总GDP,第二行是人均GDP total_gdp = gdp_ppp_rows[0].find('td', class_='infobox-data').get_text(strip=True) per_capita_gdp = gdp_ppp_rows[1].find('td', class_='infobox-data').get_text(strip=True) # 过滤参考标记和多余符号 total_gdp = total_gdp.split('[')[0].strip() per_capita_gdp = per_capita_gdp.split('[')[0].strip() print(f"GDP(PPP)总额:{total_gdp}") print(f"人均GDP(PPP):{per_capita_gdp}")
内容的提问来源于stack exchange,提问作者Mina Ashraf
相关产品推荐
相关产品推荐

