如何用Python提取HTML中标记内容并存储为数组,获取BHEL及对应数值
解决你的HTML数据提取需求
没问题,我来帮你实现从给定HTML中提取目标内容的Python代码,满足生成一维数组和单独提取股票代码及价格的需求。
实现思路
咱们用Python的BeautifulSoup库来解析HTML,这比直接字符串处理更可靠,能精准定位到各个文本节点。核心步骤如下:
- 解析HTML结构,提取标题、标的股票信息、时间信息
- 拆分出单独的股票代码
BHEL和价格80.50 - 将所有目标内容组合成符合要求的一维数组
准备工作
首先确保你安装了beautifulsoup4库,没装的话执行:
pip install beautifulsoup4
完整Python代码
from bs4 import BeautifulSoup # 你的HTML数据 html_data = '''<table width="100%"><tr><td> <div class="tphead"><h2>Option Chain (Equity Derivatives)</h2></div> </td><td align="right"> <div style="float:right; font-size:1.2em;"> <span>**Underlying Stock:** <b style="font-size:1.2em;">**BHEL** **80.50**</b> </span> <span>**As on May 11, 2018 15:30:30 IST**<a> <img onclick="refresh();" src="/live_market/resources/images/refressbtn.gif" style="cursor: pointer" title="refresh"/></a></span></div> </td></tr></table>''' # 解析HTML soup = BeautifulSoup(html_data, 'html.parser') # 提取标题 title = soup.find('h2').get_text(strip=True) # 提取标的股票相关文本并清理格式 underlying_span = soup.find('span', string=lambda text: text and 'Underlying Stock:' in text) underlying_text = underlying_span.get_text(strip=True).replace('**', '') # 拆分出独立的股票代码和价格 stock_code, stock_price = underlying_text.split(': ')[1].split() # 提取时间文本并清理格式 time_span = soup.find('span', string=lambda text: text and 'As on' in text) time_text = time_span.get_text(strip=True).replace('**', '') # 生成两种格式的目标数组:拆分版和合并版(和你给出的示例格式一致) target_array_split = [title, underlying_text, time_text] target_array_combined = [' '.join([title, underlying_text, time_text])] # 输出结果 print("目标一维数组(拆分版):", target_array_split) print("目标一维数组(合并版,匹配示例格式):", target_array_combined) print("单独提取的股票代码:", stock_code) print("单独提取的股票价格:", stock_price)
运行结果
执行代码后,你会得到:
目标一维数组(拆分版): ['Option Chain (Equity Derivatives)', 'Underlying Stock: BHEL 80.50', 'As on May 11, 2018 15:30:30 IST'] 目标一维数组(合并版,匹配示例格式): ['Option Chain (Equity Derivatives) Underlying Stock: BHEL 80.50 As on May 11, 2018 15:30:30 IST'] 单独提取的股票代码: BHEL 单独提取的股票价格: 80.50
代码里我做了两种数组格式的输出,你可以根据编码需求选择拆分版或者合并版,同时已经把BHEL和80.50单独提取成了变量,直接调用就行~
内容的提问来源于stack exchange,提问作者Marx Babu
相关产品推荐
相关产品推荐

