You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:BeautifulSoup无法捕获网页内容问题排查

Fixing Your BeautifulSoup Parsing Issues

Hey there! Let's break down why your code is failing on some sites with existing <p> tags, and fix it up with practical, reliable tweaks.

First: Fix the Critical Syntax Error

Looking at your code, there's a simple typo that's definitely causing crashes on valid sites:

source = requests.get('https://reactpodcast.com/episodes/96').textexcept:

You merged .text and except into one line—Python can't parse that! Let's rewrite the request block properly, adding better error handling for HTTP issues too:

try:
    url = 'https://reactpodcast.com/episodes/96'  # Store URL separately for later checks
    response = requests.get(url)
    response.raise_for_status()  # Catch specific HTTP errors like 404, 500, or timeouts
    source = response.text
except requests.exceptions.RequestException as e:
    print(f'Failed to fetch the site: {str(e)}')
    sys.exit()

This gives you specific error messages instead of a generic "Site does not exist"—way more helpful for debugging.

Fix the YouTube Detection Logic

Your current check for YouTube uses "https://www.youtube.com" in source, but source is the full HTML of the page, not the URL. If the page doesn't happen to include that exact string, it'll skip the YouTube logic even for valid YouTube links. Instead, check the URL you stored:

elif "youtube.com" in url:

Stop Over-Catching Exceptions

Your outer try-except wraps almost all your logic, which means any small issue (like a missing YouTube title tag) will trigger the generic "This URL is invalid" message. That's not helpful—you want to know what went wrong, not just that something did.

Instead, handle missing elements explicitly:

elif "youtube.com" in url:
    yt_title = soup.find('span', class_='watch-title')
    yt_description = soup.find('p', attrs={'id': 'eow-description'})
    if yt_title:
        print(yt_title.text.strip())
    if yt_description:
        print(yt_description.text.strip())
    if not yt_title and not yt_description:
        print("Couldn't find YouTube video details")

Clean Up Paragraph Text Extraction

When you loop through <p> tags, some might be empty or have only whitespace. Filter those out to avoid messy output, and use string joining instead of incremental concatenation (it's faster and cleaner):

allContent = '\n'.join([p.text.strip() for p in content if p.text.strip()])

Full Fixed Code

Here's the revised version with all these fixes, plus extra flexibility (like accepting a URL as a command-line argument):

import requests
import sys
from bs4 import BeautifulSoup

# Accept URL as a command-line argument for easier testing
if len(sys.argv) < 2:
    print("Please provide a URL to parse")
    sys.exit()

url = sys.argv[1]

try:
    response = requests.get(url)
    response.raise_for_status()
    source = response.text
except requests.exceptions.RequestException as e:
    print(f"Request failed: {str(e)}")
    sys.exit()

soup = BeautifulSoup(source, 'lxml')
page_title = soup.find('title')
article = soup.find('article')
paragraphs = soup.find_all('p')

# Print page title if found
if page_title:
    print(page_title.text.strip())
else:
    print("No page title found")

# Core parsing logic
if article:
    print("\n" + article.text.strip())
elif "youtube.com" in url:
    yt_title = soup.find('span', class_='watch-title')
    yt_desc = soup.find('p', attrs={'id': 'eow-description'})
    if yt_title:
        print("\n" + yt_title.text.strip())
    if yt_desc:
        print("\n" + yt_desc.text.strip())
    if not yt_title and not yt_desc:
        print("\nCould not locate YouTube video content")
elif paragraphs:
    full_content = '\n'.join([p.text.strip() for p in paragraphs if p.text.strip()])
    print("\n" + full_content)
else:
    print("\nNo extractable content found on this page")

Bonus Tips for Reliable Parsing

  • Handle dynamic content: If some sites load <p> tags with JavaScript (not in the initial HTML), BeautifulSoup can't see them. For these cases, use tools like selenium or playwright to render the page like a browser first.
  • Use specific selectors: Instead of grabbing all <p> tags, try narrowing down to containers that hold main content (e.g., soup.find('div', class_='content').find_all('p'))—this avoids pulling footer/sidebar paragraphs you don't want.
  • Test edge cases: Try your code on sites with empty <p> tags, no <article> or <p> tags, and HTTP errors to make sure it behaves gracefully.

内容的提问来源于stack exchange,提问作者Balázs Kemenesi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 10:47:32