You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python的BeautifulSoup中提取超链接内书名并去除链接标签?

解决方法:提取超链接文本并修正代码错误

问题根源

  1. book_arthur = soup.find_all("a", class_="md-crosslink") 返回的是ResultSet对象(本质是标签列表),不能像BeautifulSoup实例那样调用搜索方法,所以book_arthur("md-crosslink> center> a")会触发TypeError: 'ResultSet' object is not callable。
  2. 你要提取的是链接内的书名文本,完全不需要使用decompose()(这个方法用于删除HTML标签,和你的需求无关)。

修正后的完整代码

首先修复BeautifulSoup的拼写错误,然后直接提取每个a标签的文本内容:

import requests
from bs4 import BeautifulSoup

website = requests.get("Books.com")
soup = BeautifulSoup(website.text, "lxml")

book_links = soup.find_all("a", class_="md-crosslink")

# 遍历提取书名
for link in book_links:
    # 提取文本并去除首尾空白
    book_title = link.get_text(strip=True)
    # 可选:如果需要去掉标题末尾的数字(比如Captain Underpants1 → Captain Underpants)
    # book_title = book_title.rstrip('0123456789').strip()
    print(book_title)

运行结果

执行代码后会输出:

Harry Potter
Captain Underpants1

如果需要处理掉标题末尾的数字,取消注释代码中的字符串处理逻辑即可得到你期望的Captain Underpants。

内容的提问来源于stack exchange,提问作者Mask Shadow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 07:57:26