You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Lipidmaps网页提取非标准结构表格及SMILES数据?

解决Lipidmaps非标准表格及SMILES数据提取问题

一、提取Calculated Physicochemical Properties表格

该表格采用flex布局的div实现,无法通过标准table标签直接定位,需通过标题关联到目标容器后遍历提取:

import requests
from bs4 import BeautifulSoup

url = "https://www.lipidmaps.org/databases/lmsd/LMSL01010001"
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# 定位表格标题,找到对应的父容器
table_title = soup.find('h3', string='Calculated Physicochemical Properties')
table_container = table_title.find_next_sibling('div')

# 遍历每一行的flex项,提取属性名和对应值
physico_data = {}
for row in table_container.find_all('div', class_='flex'):
    key = row.find('div', class_='flex-grow-0 flex-shrink-0 w-60').get_text(strip=True)
    value = row.find('div', class_='flex-grow flex-shrink p-3 px-5').get_text(strip=True)
    physico_data[key] = value

# 输出提取结果
print(physico_data)

二、提取SMILES值

SMILES值在页面的属性区域,通过文本匹配定位后提取:

# 定位SMILES标签并提取对应值
smiles_label = soup.find('div', string='SMILES')
if smiles_label:
    smiles_value = smiles_label.find_next_sibling('div').get_text(strip=True)
    print("SMILES:", smiles_value)

三、关于pandas.read_table失效的说明

pandas.read_table()仅适用于读取结构化文本表格(如TSV格式),而该页面的表格是用HTML div自定义布局实现的非标准结构,pandas默认的HTML解析器无法识别这类布局,因此必须通过BeautifulSoup手动定位提取。

内容的提问来源于stack exchange,提问作者Nima Hojat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 17:40:27