You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:无法提取网页侧边栏全部按钮URL的问题(附代码)

问题分析与解决

你的代码只提取到第一个URL的原因很清晰:

  • soup.find_all('div', {'class':'side_categories'}) 获取的是整个侧边栏的容器div,页面里只有1个这样的容器,所以外层循环只会执行一次。
  • 接着在这个div里用find('a')只会返回第一个匹配的<a>标签(也就是顶部的"Books"链接),自然只能输出一个URL。

修正后的代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

url = "https://books.toscrape.com/"
genres = ["Travel", "Mystery", "Historical Fiction", "Sequential Art", "Classics", "Philosophy"]

response = requests.get(url, timeout=3)
soup = BeautifulSoup(response.content, 'html.parser')
# 定位到侧边栏里的分类列表容器
sidebar_list = soup.find('div', {'class':'side_categories'}).find('ul')
# 遍历列表里所有的<a>标签
for a_tag in sidebar_list.find_all('a'):
    genre_url = a_tag.get('href')
    print(genre_url)

输出结果

catalogue/category/books_1/index.html
catalogue/category/books/travel_2/index.html
catalogue/category/books/mystery_3/index.html
catalogue/category/books/historical-fiction_4/index.html
catalogue/category/books/sequential-art_5/index.html
catalogue/category/books/classics_6/index.html
catalogue/category/books/philosophy_7/index.html
...

如果只想提取你指定genres列表里的分类URL,可以加个文本匹配判断:

import requests
from bs4 import BeautifulSoup
import pandas as pd

url = "https://books.toscrape.com/"
genres = ["Travel", "Mystery", "Historical Fiction", "Sequential Art", "Classics", "Philosophy"]

response = requests.get(url, timeout=3)
soup = BeautifulSoup(response.content, 'html.parser')
sidebar_list = soup.find('div', {'class':'side_categories'}).find('ul')

for a_tag in sidebar_list.find_all('a'):
    # 清除标签文本里的多余空格
    genre_name = a_tag.get_text().strip()
    if genre_name in genres:
        genre_url = a_tag.get('href')
        print(genre_url)

内容的提问来源于stack exchange,提问作者Gartz man

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 22:55:16