如何用BeautifulSoup从爬取结果中提取ULTRACEMCO与7000并格式化输出?
解决方案
1. 爬取网页并获取目标H6文本
先通过requests拉取网页内容,再用BeautifulSoup解析出所有H6标签,取第一个标签的原始文本:
import requests from bs4 import BeautifulSoup url = "https://niftyinvest.com/max-pain/ULTRACEMCO?expiry=25JAN2023" resp = requests.get(url) soup = BeautifulSoup(resp.text, "html.parser") raw_h6_text = soup.find_all("h6")[0].get_text()
2. 清理冗余内容并提取目标值
针对带有多余换行、空格的文本,提供两种提取方式:
方式一:基于文本分割筛选
先清理空白字符,再通过字符串特性筛选目标值:
# 清理所有冗余空白,生成有效单词列表 cleaned_words = [word for word in raw_h6_text.split() if word.strip()] # 提取全大写的代码和数字价格 stock_code = next(w for w in cleaned_words if w.isupper()) max_pain_price = next(w for w in cleaned_words if w.isdigit()) # 按要求格式输出 print(f"{stock_code} {max_pain_price}")
方式二:基于正则表达式匹配
利用正则直接匹配目标内容,适配更灵活的文本结构:
import re # 匹配连续大写字母(代码)和连续数字(价格) match_result = re.search(r"([A-Z]+)\D*(\d+)", raw_h6_text) if match_result: print(f"{match_result.group(1)} {match_result.group(2)}")
内容的提问来源于stack exchange,提问作者Python Learner
相关产品推荐
相关产品推荐

