You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Python过滤网页爬取结果,仅提取S63.SI代码?

提取雅虎财经页面中的股票代码(S63.SI)解决方案

嘿,我来帮你搞定这个股票代码提取的问题!你现在的代码已经能抓到目标文本了,不过先得修正一个小问题:原代码里的container没有定义,应该换成page_soup,而且findAll会返回一个列表,我们用find来直接获取单个目标元素更合适。接下来给你两种可靠的提取方式:

方法一:字符串分割法

利用字符串的分割操作,直接把括号里的内容拆出来,代码简单直观:

from urllib.request import urlopen as uReq
from bs4 import BeautifulSoup as soup
import numpy as np
import pandas as pd

my_url = 'https://sg.finance.yahoo.com/quote/S63.SI/history?p=S63.SI'
uClient = uReq(my_url)
page_html = uClient.read()
uClient.close()

# html parsing
page_soup = soup(page_html, "html.parser")
# 修正:用find获取单个元素,而非findAll
target_element = page_soup.find("td", {"class":"D(ib) Fz(18px)"})
# 提取括号内的股票代码
stock_code = target_element.text.split('(')[1].split(')')[0]
print(stock_code)  # 输出:S63.SI

方法二:正则表达式法

如果以后遇到更复杂的文本格式,正则表达式会更灵活,能精准匹配括号内的内容:

from urllib.request import urlopen as uReq
from bs4 import BeautifulSoup as soup
import numpy as np
import pandas as pd
import re  # 导入正则模块

my_url = 'https://sg.finance.yahoo.com/quote/S63.SI/history?p=S63.SI'
uClient = uReq(my_url)
page_html = uClient.read()
uClient.close()

# html parsing
page_soup = soup(page_html, "html.parser")
target_element = page_soup.find("td", {"class":"D(ib) Fz(18px)"})
# 匹配括号内的任意内容
pattern = r'\((.*?)\)'
match_result = re.search(pattern, target_element.text)
if match_result:
    stock_code = match_result.group(1)
    print(stock_code)  # 输出:S63.SI

两种方法都能完美提取出你想要的S63.SI,你可以根据自己的习惯选择~

内容的提问来源于stack exchange,提问作者Lim Han Yang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 14:33:14