如何使用BeautifulSoup从指定p标签中提取Acid与Alcohol的数值?
提取特定p标签的数值方案
可以通过过滤文本内容精准定位包含Acid和Alcohol的p标签,而非获取所有p标签,下面是两种实用方法:
方法1:用lambda表达式筛选文本
借助find_all的text参数配合lambda,只保留含目标关键词的p标签:
from bs4 import BeautifulSoup # 替换为你的实际HTML内容 html_content = """ <p>Acid(5.9 g/L)</p> <p>Alcohol(14.5%)</p> <p>其他无关内容</p> """ soup = BeautifulSoup(html_content, 'html.parser') # 筛选包含Acid或Alcohol的p标签 target_elements = soup.find_all('p', text=lambda t: t and ('Acid' in t or 'Alcohol' in t)) # 提取数值并整理成字典 extracted_values = {} for elem in target_elements: clean_text = elem.get_text(strip=True) if 'Acid' in clean_text: extracted_values['Acid'] = clean_text.split('(')[1].rstrip(')') elif 'Alcohol' in clean_text: extracted_values['Alcohol'] = clean_text.split('(')[1].rstrip(')') print(extracted_values) # 输出:{'Acid': '5.9 g/L', 'Alcohol': '14.5%'}
方法2:用正则表达式匹配
如果需要更灵活的文本匹配规则,可通过正则指定关键词:
import re from bs4 import BeautifulSoup soup = BeautifulSoup(html_content, 'html.parser') # 正则匹配包含Acid或Alcohol的文本 keyword_pattern = re.compile(r'(Acid|Alcohol)') target_elements = soup.find_all('p', text=keyword_pattern) # 提取数值 extracted_values = {} for elem in target_elements: clean_text = elem.get_text(strip=True) key = 'Acid' if 'Acid' in clean_text else 'Alcohol' extracted_values[key] = clean_text.split('(')[1].rstrip(')') print(extracted_values)
补充说明
- 若目标p标签有专属class/id属性,可直接用
find_all('p', class_='目标类名')缩小范围,提升效率; - 若数值格式固定,也能直接用正则提取数值部分,比如
re.search(r'\d+\.?\d*\s?%?g/L?', clean_text).group()。
内容的提问来源于stack exchange,提问作者DJ-coding
相关产品推荐
相关产品推荐

