如何使用Pyquery的.attr()提取网页指定URL?
提取目标URL的实现方法
方法一:使用.attr()结合属性选择器直接定位
你可以通过属性选择器精准匹配目标<a>标签,再用.attr('href')提取链接。目标URL的特征是包含historical-mortgage-origination-estimates.xlsx?sfvrsn=8c6933cb_5,用包含匹配选择器[href*="xxx"]即可定位:
from pyquery import PyQuery as pq import requests url = "https://www.mba.org/news-and-research/forecasts-and-commentary" content = requests.get(url).content doc = pq(content) # 定位目标div下所有符合条件的a标签 target_links = doc("#ContentPlaceholder_C012_Col01 a[href*='historical-mortgage-origination-estimates.xlsx?sfvrsn=8c6933cb_5']") # 提取第一个匹配的链接 first_target_href = target_links.eq(0).attr('href') print(first_target_href) # 提取所有符合条件的链接(遍历输出) for link in target_links.items(): print(link.attr('href'))
方法二:遍历所有a标签筛选目标URL
如果需要更灵活的筛选逻辑,可以先获取目标区域内的所有<a>标签,再逐个判断href是否符合要求:
from pyquery import PyQuery as pq import requests url = "https://www.mba.org/news-and-research/forecasts-and-commentary" content = requests.get(url).content doc = pq(content) target_div = doc("#ContentPlaceholder_C012_Col01") # 获取div下所有a标签 all_links = target_div.find('a') # 筛选并收集目标URL target_hrefs = [] for link in all_links.items(): href = link.attr('href') if href and 'historical-mortgage-origination-estimates.xlsx?sfvrsn=8c6933cb_5' in href: target_hrefs.append(href) print(target_hrefs)
补充说明
.attr('href')会返回元素的href属性值,要是得到的是相对路径,直接拼接网站域名https://www.mba.org就能生成完整可访问的URL。- 两种方法各有适用场景:属性选择器定位效率更高,遍历筛选则适合需要额外逻辑判断的情况。
内容的提问来源于stack exchange,提问作者prashanth manohar
相关产品推荐
相关产品推荐

