You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫用BeautifulSoup提取作者名与处理价格空格异常问题

爬取书籍信息代码修复方案

配套页面结构截图:
页面结构截图

问题修复说明

  • 第一个问题是变量引用错误:判断作者标签存在时,你将提取到的纯文本存入了author变量,但最终写入字典时仍调用了原始标签对象author_check,因此会返回完整标签内容,修正引用逻辑即可。
  • 第二个问题是不间断空格转换:直接对提取到的价格字符串调用replace('\xa0', ' ')方法,即可将不间断空格替换为普通空格。
  • 额外优化:删除无效的from pip._internal.network.utils import HEADERS导入语句,避免和自定义的HEADERS变量产生冲突。

修复后完整代码

import requests
from bs4 import BeautifulSoup

URL = "https://www.moscowbooks.ru/books/?sortby=name&sortdown=false"
HEADERS = {"user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_6) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/14.0.3 Safari/605.1.15", "accept" : "*/*"}
HOST = "https://www.moscowbooks.ru"

def get_html(url, params=None):
    r = requests.get(url, headers = HEADERS, params = params)
    return r


def get_content(html):
    soup = BeautifulSoup(html, "html.parser")
    items = soup.find_all("div", class_ = "catalog__item")
    books = []
    for item in items:
        author_check = item.find("a", class_="author-name")
        if author_check:
            # 提取纯文本并去除首尾空白
            author = author_check.get_text(strip=True)
        else:
            author = "Автор не указан"
        # 处理价格中的不间断空格
        cost = item.find("div", class_="book-preview__price").get_text(strip=True).replace('\xa0', ' ')
        books.append({
            "title": item.find("div", class_ = "book-preview__title").get_text(strip=True),
            "author": author,
            "link": HOST + item.find("a", class_ = "book-preview__title-link").get("href"),
            "cost": cost,
        })
    print(books)
    print(len(books))

def parse():
    html = get_html(URL)
    if html.status_code == 200:
        get_content(html.text)
    else:
        print("Error")

parse()

内容的提问来源于stack exchange,提问作者CMDR_Mark

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 00:15:03