You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HTML表格ISIN代码抓取遇阻:现有方法报错,求稳定方案

解决ISIN代码抓取时的AttributeError问题

你的问题核心在于原方法过度依赖find(text=search)定位元素——一旦页面结构微调(比如文本前后多了空格、标签嵌套逻辑变化),这个方法就会返回None,后续调用parent自然触发AttributeError。下面给你几个更稳定的解决方案:

方法1:通过标签对应关系精准定位

利用表格中th(标题)和td(内容)的绑定关系,先找到包含"ISIN:"的th,再直接获取其相邻的td内容:

from urllib.request import urlopen
from bs4 import BeautifulSoup

quote_page = 'https://www.hl.co.uk/shares/shares-search-results/t/tesco-6-2029'
page = urlopen(quote_page)
soup = BeautifulSoup(page, 'html.parser')

# 用lambda匹配文本,允许前后有空格,避免严格匹配失败
isin_th = soup.find('th', text=lambda text: text and 'ISIN:' in text.strip())
if isin_th:
    # 获取相邻的td标签内容,自动处理空格
    bond_code = isin_th.find_next_sibling('td').get_text(strip=True)
    print(bond_code)
else:
    print("未找到ISIN信息")

方法2:CSS选择器简化写法(需配合lxml解析器)

如果安装了lxml解析器(执行pip install lxml即可),可以用更简洁的CSS选择器直接定位目标:

from urllib.request import urlopen
from bs4 import BeautifulSoup

quote_page = 'https://www.hl.co.uk/shares/shares-search-results/t/tesco-6-2029'
page = urlopen(quote_page)
# 换成lxml解析器,支持更多CSS选择器语法
soup = BeautifulSoup(page, 'lxml')

bond_code_elem = soup.select_one('tr.table-alt th:contains("ISIN:") + td')
if bond_code_elem:
    bond_code = bond_code_elem.get_text(strip=True)
    print(bond_code)
else:
    print("未找到ISIN信息")

方法3:遍历表格行(兼容性最强)

如果页面结构变动较大,还可以遍历所有表格行,逐一检查是否包含ISIN信息:

from urllib.request import urlopen
from bs4 import BeautifulSoup

quote_page = 'https://www.hl.co.uk/shares/shares-search-results/t/tesco-6-2029'
page = urlopen(quote_page)
soup = BeautifulSoup(page, 'html.parser')

# 先定位目标表格
target_table = soup.find('table', class_='factsheet-table table-no-border spacer-bottom')
if target_table:
    # 遍历所有行
    for row in target_table.find_all('tr'):
        row_title = row.find('th')
        if row_title and 'ISIN:' in row_title.get_text(strip=True):
            row_content = row.find('td')
            if row_content:
                bond_code = row_content.get_text(strip=True)
                print(bond_code)
                break
    else:
        print("未找到ISIN信息")
else:
    print("未找到目标表格")

关键优化要点

  • 增加空值判断:每次调用find/select_one后,先检查返回值是否为None,再执行后续操作,避免报错
  • 用get_text(strip=True)替代contents[0].strip():get_text会自动处理标签内的所有文本,避免因嵌套标签导致的内容获取失败
  • 灵活匹配文本:允许文本前后有空格,避免页面微小调整导致匹配失效

内容的提问来源于stack exchange,提问作者Charlie Auckland

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:19:07