如何用BeautifulSoup提取论坛链接?现有代码无法获取论坛超链接
解决方法
你的代码只提取以https://开头的绝对链接,但论坛里的大量内部链接是相对路径(比如/threads/xxx)或其他格式,所以会遗漏。可以通过以下几种方式修复:
方案1:捕获所有非空链接,统一转为绝对路径
直接提取所有带href属性的<a>标签,再用urllib.parse.urljoin把相对路径转换成绝对路径,确保能拿到完整链接:
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin url = "https://www.diabetesdaily.com/forum/forums/type-1-diabetes.9/" r1b = requests.get(url) if r1b.status_code == 200: print("Accessible.") sp1b = BeautifulSoup(r1b.text, 'lxml') for link in sp1b.find_all('a', href=True): # 把相对路径转为绝对路径 full_url = urljoin(url, link.get('href')) print(full_url)
方案2:调整正则匹配更多链接格式
修改正则表达式,覆盖HTTPS/HTTP、协议相对路径(//xxx)、根相对路径(/xxx):
import requests from bs4 import BeautifulSoup import re url = "https://www.diabetesdaily.com/forum/forums/type-1-diabetes.9/" r1b = requests.get(url) if r1b.status_code == 200: print("Accessible.") sp1b = BeautifulSoup(r1b.text, 'lxml') # 匹配http/https开头、//开头、/开头的链接 pattern = re.compile(r'^(https?://|//|/)') for link in sp1b.find_all('a', attrs={'href': pattern}): # 对//开头的链接补全协议 href = link.get('href') if href.startswith('//'): href = 'https:' + href # 对/开头的链接补全域名 elif href.startswith('/'): href = 'https://www.diabetesdaily.com' + href print(href)
方案3:处理动态加载链接(如果适用)
如果论坛的部分内容是通过JavaScript动态渲染的,requests无法获取到这部分链接,此时需要用浏览器渲染工具(比如Selenium):
from selenium import webdriver from selenium.webdriver.common.by import By from urllib.parse import urljoin url = "https://www.diabetesdaily.com/forum/forums/type-1-diabetes.9/" driver = webdriver.Chrome() # 需要提前安装ChromeDriver driver.get(url) # 等待页面加载完成(可根据实际情况调整等待时间或用显式等待) driver.implicitly_wait(10) for link in driver.find_elements(By.TAG_NAME, 'a'): href = link.get_attribute('href') if href: full_url = urljoin(url, href) print(full_url) driver.quit()
内容的提问来源于stack exchange,提问作者KiuSandy
相关产品推荐
相关产品推荐

