Python提取HTML中href链接返回None,求适配所有标签的通用方案
问题分析
你通过soup.select('.titleline')获取的.titleline元素本身不具备href属性,href是它内部<a>标签的属性,所以直接调用.get('href')会返回None。修改后的代码通过.find('a')找到子标签获取href,但你希望代码能适配任意包含href的标签(不管是当前元素本身,还是它的后代元素)。
解决方案
以下是几种通用的href提取实现方式:
方法1:提取元素及其所有后代中带href的链接
使用BeautifulSoup的find_all方法,筛选所有带有href属性的元素,无需限定标签类型:
import requests from bs4 import BeautifulSoup res = requests.get('https://news.ycombinator.com/') soup = BeautifulSoup(res.text, 'html.parser') links = soup.select('.titleline') votes = soup.select('.score') def create_custom_hn(links, votes): hn = [] for item in links: # 查找当前元素及其所有后代中,所有带href属性的元素 href_elements = item.find_all(attrs={'href': True}) for elem in href_elements: href = elem.get('href') if href: # 过滤空链接 hn.append(href) return hn print(create_custom_hn(links, votes))
方法2:用CSS选择器直接定位目标范围内的带href元素
不需要单独遍历.titleline元素,直接用CSS选择器.titleline [href]获取所有符合条件的元素,简化代码:
import requests from bs4 import BeautifulSoup res = requests.get('https://news.ycombinator.com/') soup = BeautifulSoup(res.text, 'html.parser') # 直接获取.titleline下所有带href属性的元素 href_elements = soup.select('.titleline [href]') votes = soup.select('.score') def create_custom_hn(href_elements, votes): hn = [elem.get('href') for elem in href_elements if elem.get('href')] return hn print(create_custom_hn(href_elements, votes))
方法3:通用单元素href提取函数(含去重)
如果需要同时处理标题和链接,可编写辅助函数提取单个元素内的所有有效href,还支持去重:
import requests from bs4 import BeautifulSoup res = requests.get('https://news.ycombinator.com/') soup = BeautifulSoup(res.text, 'html.parser') links = soup.select('.titleline') votes = soup.select('.score') def extract_hrefs(element): """提取单个元素及其后代中所有有效的href链接(去重)""" hrefs = [] # 先检查元素自身是否有href self_href = element.get('href') if self_href: hrefs.append(self_href) # 再遍历后代元素的href for child in element.find_all(attrs={'href': True}): href = child.get('href') if href and href not in hrefs: hrefs.append(href) return hrefs def create_custom_hn(links, votes): hn = [] for item in links: title = item.getText(strip=True) hrefs = extract_hrefs(item) hn.append({ 'title': title, 'hrefs': hrefs }) return hn print(create_custom_hn(links, votes))
说明
attrs={'href': True}用于匹配所有带有href属性的元素,支持<a>、<link>等任意标签。if href判断可过滤href=""这类空链接,避免无效数据。- 方法3中的去重逻辑可根据实际需求选择是否保留。
内容的提问来源于stack exchange,提问作者Auprell Edwards
相关产品推荐
相关产品推荐

