BeautifulSoup提取href以peak.aspx?pid=10开头的特定链接如何实现?
实现方案
你不需要做字符串分割操作,两种实现方式都可以,优先推荐第一种更简洁的方案:
方案1:直接使用BeautifulSoup内置的CSS属性前缀选择器
CSS选择器支持[attr^=value]语法,用于匹配属性值以指定字符串开头的元素,不需要额外写循环遍历逻辑,直接一步筛选出符合要求的a标签:
from urllib.request import urlopen from bs4 import BeautifulSoup import pandas as pd url = 'https://www.peakbagger.com/list.aspx?lid=5651' html = urlopen(url) soup = BeautifulSoup(html, 'html.parser') # 直接筛选href以peak.aspx?pid=10开头的a标签 target_a = soup.select('a[href^="peak.aspx?pid=10"]') print(target_a) # 如果要单独提取href值可以加以下语句 href_list = [a['href'] for a in target_a] print(href_list)
方案2:循环遍历已有a标签筛选
如果你已经拿到了所有a标签的列表,可以通过字符串内置的startswith()方法判断过滤,不需要做字符串分割:
# 你之前拿到的所有a标签 a_list = soup.select("a:nth-of-type(1)") target_a = [] for a in a_list: # 先获取href属性,没有href的a标签默认给空字符串避免报错 href = a.get('href', '') if href.startswith('peak.aspx?pid=10'): target_a.append(a)
内容的提问来源于stack exchange,提问作者Lemon
相关产品推荐
相关产品推荐

