如何用BeautifulSoup提取HTML表格中苹果和香蕉的百分比数据
解决BeautifulSoup提取指定水果对应百分比的问题
你可以通过两种方式实现仅提取苹果和香蕉的对应百分比并按指定格式输出:
方法一:遍历表格行筛选目标水果
这种方式灵活性更强,方便后续扩展提取更多水果:
from bs4 import BeautifulSoup HTMLFile = open("stock.html", "r") index = HTMLFile.read() soup = BeautifulSoup(index, 'html.parser') # 指定需要提取的水果标签 target_fruits = ["apples:", "bananas:"] # 遍历表格的所有行 for row in soup.select(".stock tr"): # 获取当前行的水果名称单元格 fruit_label = row.find("th", scope="row") if fruit_label and fruit_label.text.strip() in target_fruits: # 提取该行的百分比值 percentage = row.find("td", class_="high").text.strip() # 按要求格式输出,去掉原标签文本末尾的冒号 print(f"{fruit_label.text.strip().rstrip(':')}: {percentage}")
方法二:使用CSS选择器直接定位
如果目标水果固定,也可以用更简洁的CSS选择器直接定位到对应元素:
from bs4 import BeautifulSoup HTMLFile = open("stock.html", "r") index = HTMLFile.read() soup = BeautifulSoup(index, 'html.parser') # 定位苹果对应的百分比 apple_percent = soup.select_one("tr:has(th:contains('apples:')) .high").text.strip() # 定位香蕉对应的百分比 banana_percent = soup.select_one("tr:has(th:contains('bananas:')) .high").text.strip() # 按指定格式输出 print(f"apples: {apple_percent}") print(f"bananas: {banana_percent}")
说明
- 方法一通过遍历行并筛选目标水果,适合后续需要提取更多品类的场景;
- 方法二利用
tr:has()选择器直接匹配包含指定水果名称的行,代码更简洁; - 两种方法最终都会输出:
apples: 98.2% bananas: 94.2%
内容的提问来源于stack exchange,提问作者cosmos
相关产品推荐
相关产品推荐

