You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取HTML中href链接返回None,求适配所有标签的通用方案

问题分析

你通过soup.select('.titleline')获取的.titleline元素本身不具备href属性,href是它内部<a>标签的属性,所以直接调用.get('href')会返回None。修改后的代码通过.find('a')找到子标签获取href,但你希望代码能适配任意包含href的标签(不管是当前元素本身,还是它的后代元素)。

解决方案

以下是几种通用的href提取实现方式:

方法1:提取元素及其所有后代中带href的链接

使用BeautifulSoup的find_all方法,筛选所有带有href属性的元素,无需限定标签类型:

import requests 
from bs4 import BeautifulSoup

res = requests.get('https://news.ycombinator.com/')
soup = BeautifulSoup(res.text, 'html.parser')
links = soup.select('.titleline')
votes = soup.select('.score')


def create_custom_hn(links, votes):
    hn = []
    for item in links:
        # 查找当前元素及其所有后代中,所有带href属性的元素
        href_elements = item.find_all(attrs={'href': True})
        for elem in href_elements:
            href = elem.get('href')
            if href:  # 过滤空链接
                hn.append(href)
    return hn

print(create_custom_hn(links, votes))

方法2:用CSS选择器直接定位目标范围内的带href元素

不需要单独遍历.titleline元素,直接用CSS选择器.titleline [href]获取所有符合条件的元素,简化代码:

import requests 
from bs4 import BeautifulSoup

res = requests.get('https://news.ycombinator.com/')
soup = BeautifulSoup(res.text, 'html.parser')
# 直接获取.titleline下所有带href属性的元素
href_elements = soup.select('.titleline [href]')
votes = soup.select('.score')


def create_custom_hn(href_elements, votes):
    hn = [elem.get('href') for elem in href_elements if elem.get('href')]
    return hn

print(create_custom_hn(href_elements, votes))

方法3:通用单元素href提取函数(含去重)

如果需要同时处理标题和链接,可编写辅助函数提取单个元素内的所有有效href,还支持去重:

import requests 
from bs4 import BeautifulSoup

res = requests.get('https://news.ycombinator.com/')
soup = BeautifulSoup(res.text, 'html.parser')
links = soup.select('.titleline')
votes = soup.select('.score')

def extract_hrefs(element):
    """提取单个元素及其后代中所有有效的href链接(去重)"""
    hrefs = []
    # 先检查元素自身是否有href
    self_href = element.get('href')
    if self_href:
        hrefs.append(self_href)
    # 再遍历后代元素的href
    for child in element.find_all(attrs={'href': True}):
        href = child.get('href')
        if href and href not in hrefs:
            hrefs.append(href)
    return hrefs

def create_custom_hn(links, votes):
    hn = []
    for item in links:
        title = item.getText(strip=True)
        hrefs = extract_hrefs(item)
        hn.append({
            'title': title,
            'hrefs': hrefs
        })
    return hn

print(create_custom_hn(links, votes))
说明
  • attrs={'href': True}用于匹配所有带有href属性的元素,支持<a>、<link>等任意标签。
  • if href判断可过滤href=""这类空链接,避免无效数据。
  • 方法3中的去重逻辑可根据实际需求选择是否保留。

内容的提问来源于stack exchange,提问作者Auprell Edwards

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 22:13:41