You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取无结果问题求助:爬取Books to Scrape网站返回空DataFrame

Fixing Your Empty Scraping Result Issue

Looking at your code, there are a few key issues causing the empty DataFrame. Let's break them down and fix them step by step:

Your get_links function contains nested functions that are never called, and it doesn't return any value (so it defaults to None). When you pass None to extract_info, the loop over links doesn't run, leaving your all_books list empty.

2. Incorrect URL Construction

String concatenation for building book URLs leads to missing slashes (e.g., http://books.toscrape.com/cataloguea-light-in-the-attic_1000/index.html instead of the correct path with a slash after catalogue). Using urljoin from the urllib.parse module fixes this by properly handling relative paths.

3. Unreliable Price Extraction

While your current price selector works, targeting the specific price_color class makes the code more robust in case the page structure changes slightly.


Corrected Code

Here's the fixed version of your code with explanations:

import pandas as pd
import requests
from bs4 import BeautifulSoup as bs
from urllib.parse import urljoin  # Import for safe URL construction

all_books = []

def get_page(url):
    page = requests.get(url)
    status = page.status_code
    soup = bs(page.text, "lxml")
    return [soup, status]

def get_links(soup, base_url):
    links = []
    listings = soup.find_all(class_="product_pod")
    for listing in listings:
        # Extract relative link from the product pod
        relative_link = listing.find("h5").a.get("href")
        # Build full absolute URL using urljoin to avoid path errors
        full_link = urljoin(base_url, relative_link)
        links.append(full_link)
    return links  # Return the list of book links

def extract_info(links):
    for link in links:
        res = requests.get(link).text
        book_soup = bs(res, "lxml")
        # Extract title from the h1 tag in product_main
        title = book_soup.find(class_="col-sm-6 product_main").h1.text.strip()
        # Extract price using the specific price_color class for reliability
        price = book_soup.find(class_="price_color").text.strip()
        book = {"title": title, "price": price}
        all_books.append(book)

pg = 1
while True:
    url = f"http://books.toscrape.com/catalogue/page-{pg}.html"
    soup_status = get_page(url)
    if soup_status[1] == 200:
        print(f"scraping page {pg}")
        # Get valid book links from the current page
        book_links = get_links(soup_status[0], url)
        # Process each link to extract book info
        extract_info(book_links)
        pg += 1
    else:
        print("The End")
        break

df = pd.DataFrame(all_books)
print(df)

Key Changes Made:

  • Removed unused nested functions in get_links and ensured it returns the list of valid book links.
  • Added urljoin to safely construct absolute URLs from relative paths, eliminating broken links.
  • Updated price extraction to target the price_color class directly for better reliability.
  • Ensured valid links are passed to extract_info, so the loop runs and populates all_books.

When you run this code, it should correctly scrape all books from the site and populate your DataFrame with titles and prices.

内容的提问来源于stack exchange,提问作者BJ Creative

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 14:24:07