You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup 4解析HTML时for循环无法运行求助

问题:无法用Beautiful Soup提取Etsy主页的所有链接

我正在参考Beautiful Soup文档学习用法,对Python不太熟悉,可能没发现语法错误。以下代码本该输出Etsy主页的所有链接,但没生效。文档有类似示例,可能漏了要点。仅打印HTML内容正常,但for循环无法工作。

#!/usr/bin/python3

# import library
from bs4 import BeautifulSoup
import requests
import os.path
from os import path

# Request to website and download HTML contents
url='https://www.etsy.com/?utm_source=google&utm_medium=cpc&utm_term=etsy_e&utm_campaign=Search_US_Brand_GGL_ENG_General-Brand_Core_All_Exact&utm_ag=A1&utm_custom1=_k_Cj0KCQiAi8KfBhCuARIsADp-A54MzODz8nRIxO2LnGcB8Ezc3_q40IQk9HygcSzz9fPmPWnrITz8InQaAt5oEALw_wcB_k_&utm_content=go_227553629_16342445429_536666953103_kwd-1818581752_c_&utm_custom2=227553629&gclid=Cj0KCQiAi8KfBhCuARIsADp-A54MzODz8nRIxO2LnGcB8Ezc3_q40IQk9HygcSzz9fPmPWnrITz8InQaAt5oEALw_wcB'
req=requests.get(url)
content=req.text

soup=BeautifulSoup(content, 'html.parser')

for x in soup.head.find_all('a'):
    print(x.get('href'))

问题原因与解决方法

核心问题

你写的soup.head.find_all('a')只在HTML的<head>标签范围内查找链接,但网页的绝大多数可点击链接都在<body>标签里,<head>里一般只有少量关联样式、脚本的链接(甚至可能没有),所以循环自然没有输出内容。

修复方案

  1. 扩大查找范围:把查找对象从soup.head改成整个文档soup,或者明确指定soup.body,这样就能遍历页面中所有的<a>标签。
  2. 过滤无效链接:可以在查找时加上href=True的条件,直接过滤掉没有href属性的<a>标签,避免输出None。
  3. 清理冗余代码:你导入的os.path相关模块没用到,可以删掉,减少代码冗余。

修改后的代码

#!/usr/bin/python3

from bs4 import BeautifulSoup
import requests

# 请求网页并获取HTML内容
url='https://www.etsy.com/?utm_source=google&amp;utm_medium=cpc&amp;utm_term=etsy_e&amp;utm_campaign=Search_US_Brand_GGL_ENG_General-Brand_Core_All_Exact&amp;utm_ag=A1&amp;utm_custom1=_k_Cj0KCQiAi8KfBhCuARIsADp-A54MzODz8nRIxO2LnGcB8Ezc3_q40IQk9HygcSzz9fPmPWnrITz8InQaAt5oEALw_wcB_k_&amp;utm_content=go_227553629_16342445429_536666953103_kwd-1818581752_c_&amp;utm_custom2=227553629&amp;gclid=Cj0KCQiAi8KfBhCuARIsADp-A54MzODz8nRIxO2LnGcB8Ezc3_q40IQk9HygcSzz9fPmPWnrITz8InQaAt5oEALw_wcB'
req=requests.get(url)
content=req.text

soup=BeautifulSoup(content, 'html.parser')

# 遍历整个页面中带有href属性的a标签
for x in soup.find_all('a', href=True):
    print(x.get('href'))

内容的提问来源于stack exchange,提问作者Nathan Brannan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 09:56:20