Python爬虫用BeautifulSoup提取作者名与处理价格空格异常问题
爬取书籍信息代码修复方案
配套页面结构截图:
问题修复说明
- 第一个问题是变量引用错误:判断作者标签存在时,你将提取到的纯文本存入了
author变量,但最终写入字典时仍调用了原始标签对象author_check,因此会返回完整标签内容,修正引用逻辑即可。 - 第二个问题是不间断空格转换:直接对提取到的价格字符串调用
replace('\xa0', ' ')方法,即可将不间断空格替换为普通空格。 - 额外优化:删除无效的
from pip._internal.network.utils import HEADERS导入语句,避免和自定义的HEADERS变量产生冲突。
修复后完整代码
import requests from bs4 import BeautifulSoup URL = "https://www.moscowbooks.ru/books/?sortby=name&sortdown=false" HEADERS = {"user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_6) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/14.0.3 Safari/605.1.15", "accept" : "*/*"} HOST = "https://www.moscowbooks.ru" def get_html(url, params=None): r = requests.get(url, headers = HEADERS, params = params) return r def get_content(html): soup = BeautifulSoup(html, "html.parser") items = soup.find_all("div", class_ = "catalog__item") books = [] for item in items: author_check = item.find("a", class_="author-name") if author_check: # 提取纯文本并去除首尾空白 author = author_check.get_text(strip=True) else: author = "Автор не указан" # 处理价格中的不间断空格 cost = item.find("div", class_="book-preview__price").get_text(strip=True).replace('\xa0', ' ') books.append({ "title": item.find("div", class_ = "book-preview__title").get_text(strip=True), "author": author, "link": HOST + item.find("a", class_ = "book-preview__title-link").get("href"), "cost": cost, }) print(books) print(len(books)) def parse(): html = get_html(URL) if html.status_code == 200: get_content(html.text) else: print("Error") parse()
内容的提问来源于stack exchange,提问作者CMDR_Mark
相关产品推荐
相关产品推荐

