如何借助Node.js库轻松获取网站的Reader View内容?
Hey there! I totally feel your pain—Safari's Reader View is such a game-changer for cutting through ads, sidebars, and all the other noise to get clean, focused content (text + images). Let me share some tried-and-true methods to replicate that behavior in your scraping projects:
1. Use Purpose-Built Content Extraction Libraries
These tools are designed specifically to identify and pull article-like content from web pages, just like Reader View does:
newspaper3k (Python)
This is one of the most popular options for article extraction. It automatically detects the main body text, images, and even metadata.
pip install newspaper3k
Example usage:
from newspaper import Article url = "your-target-url-here" article = Article(url) article.download() article.parse() # Get clean main text print(article.text) # Get list of image URLs from the article print(article.images)
trafilatura (Python)
Another excellent library focused on high-quality text extraction. It’s great at filtering out irrelevant content and preserving structure.
pip install trafilatura
Example usage:
import trafilatura url = "your-target-url-here" downloaded = trafilatura.fetch_url(url) # Extract text and images (images are included in the output if you use the right parameters) result = trafilatura.extract(downloaded, include_images=True) print(result)
2. Browser Automation with Native Reader Mode
If you want to mirror Safari’s exact Reader View logic, you can use browser automation tools to trigger the browser’s built-in reader mode, then scrape the cleaned content directly.
Playwright (Cross-browser support)
Playwright works with Chrome, Firefox, and Safari, and lets you enable reader mode programmatically:
pip install playwright playwright install
Example usage:
from playwright.sync_api import sync_playwright with sync_playwright() as p: browser = p.webkit.launch(headless=False) # Use WebKit for Safari-like engine page = browser.new_page() page.goto("your-target-url-here") # Trigger Safari-style reader mode page.evaluate(""" if (document.body.classList.contains('reader-mode')) { document.body.classList.remove('reader-mode'); } else { document.body.classList.add('reader-mode'); } """) # Wait for reader view to load, then extract content reader_content = page.locator(".reader-content").inner_text() reader_images = [img.get_attribute("src") for img in page.locator(".reader-content img").all()] print(reader_content) print(reader_images) browser.close()
3. Manual Semantic HTML Parsing (For Custom Cases)
If you’re dealing with a specific set of websites, you can leverage semantic HTML tags that most sites use for main content:
- Look for
<article>or<main>tags—these usually wrap the primary content - Target common class names like
.article-body,.post-content, or.main-content - Extract
<p>tags for text and<img>tags withsrcattributes for images
This approach requires more site-specific tweaking but gives you full control.
Quick Example with BeautifulSoup:
from bs4 import BeautifulSoup import requests url = "your-target-url-here" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # Try to find the main content container main_content = soup.find("article") or soup.find("main") or soup.find(class_="article-body") if main_content: # Extract text from all <p> tags text = "\n".join([p.get_text() for p in main_content.find_all("p")]) # Extract image URLs images = [img["src"] for img in main_content.find_all("img")] print(text) print(images)
Final Recommendation
Start with newspaper3k or trafilatura—they’re fast, require minimal code, and work for most sites. If you need to match Safari’s Reader View exactly, go with the browser automation route using WebKit (Playwright's Safari engine) to get the closest behavior.
内容的提问来源于stack exchange,提问作者Dominik

