You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何借助Node.js库轻松获取网站的Reader View内容?

Replicate iOS Safari Reader View for Web Scraping

Hey there! I totally feel your pain—Safari's Reader View is such a game-changer for cutting through ads, sidebars, and all the other noise to get clean, focused content (text + images). Let me share some tried-and-true methods to replicate that behavior in your scraping projects:

1. Use Purpose-Built Content Extraction Libraries

These tools are designed specifically to identify and pull article-like content from web pages, just like Reader View does:

newspaper3k (Python)

This is one of the most popular options for article extraction. It automatically detects the main body text, images, and even metadata.

pip install newspaper3k

Example usage:

from newspaper import Article

url = "your-target-url-here"
article = Article(url)
article.download()
article.parse()

# Get clean main text
print(article.text)
# Get list of image URLs from the article
print(article.images)

trafilatura (Python)

Another excellent library focused on high-quality text extraction. It’s great at filtering out irrelevant content and preserving structure.

pip install trafilatura

Example usage:

import trafilatura

url = "your-target-url-here"
downloaded = trafilatura.fetch_url(url)
# Extract text and images (images are included in the output if you use the right parameters)
result = trafilatura.extract(downloaded, include_images=True)

print(result)

2. Browser Automation with Native Reader Mode

If you want to mirror Safari’s exact Reader View logic, you can use browser automation tools to trigger the browser’s built-in reader mode, then scrape the cleaned content directly.

Playwright (Cross-browser support)

Playwright works with Chrome, Firefox, and Safari, and lets you enable reader mode programmatically:

pip install playwright
playwright install

Example usage:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.webkit.launch(headless=False)  # Use WebKit for Safari-like engine
    page = browser.new_page()
    page.goto("your-target-url-here")
    
    # Trigger Safari-style reader mode
    page.evaluate("""
        if (document.body.classList.contains('reader-mode')) {
            document.body.classList.remove('reader-mode');
        } else {
            document.body.classList.add('reader-mode');
        }
    """)
    
    # Wait for reader view to load, then extract content
    reader_content = page.locator(".reader-content").inner_text()
    reader_images = [img.get_attribute("src") for img in page.locator(".reader-content img").all()]
    
    print(reader_content)
    print(reader_images)
    
    browser.close()

3. Manual Semantic HTML Parsing (For Custom Cases)

If you’re dealing with a specific set of websites, you can leverage semantic HTML tags that most sites use for main content:

  • Look for <article> or <main> tags—these usually wrap the primary content
  • Target common class names like .article-body, .post-content, or .main-content
  • Extract <p> tags for text and <img> tags with src attributes for images

This approach requires more site-specific tweaking but gives you full control.

Quick Example with BeautifulSoup:

from bs4 import BeautifulSoup
import requests

url = "your-target-url-here"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")

# Try to find the main content container
main_content = soup.find("article") or soup.find("main") or soup.find(class_="article-body")

if main_content:
    # Extract text from all <p> tags
    text = "\n".join([p.get_text() for p in main_content.find_all("p")])
    # Extract image URLs
    images = [img["src"] for img in main_content.find_all("img")]
    print(text)
    print(images)

Final Recommendation

Start with newspaper3k or trafilatura—they’re fast, require minimal code, and work for most sites. If you need to match Safari’s Reader View exactly, go with the browser automation route using WebKit (Playwright's Safari engine) to get the closest behavior.

内容的提问来源于stack exchange,提问作者Dominik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:50:41