You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫爬取ASCO会议页面遇内容拦截问题求助

Fixing "Blocked Content" When Scraping ASCO Meeting Library

Hey, I’ve run into similar anti-scraping blocks with Angular-powered sites before, so let’s break down what’s happening and how to fix it:

  • ASCO’s site uses Angular (client-side rendering), meaning the div.inner content you’re targeting isn’t present in the initial raw HTML response—it gets rendered by JavaScript after the page loads. Your current urllib code only fetches the static base HTML, which doesn’t include the meeting data you want.
  • The site is also flagging your request as a bot because generic urllib requests lack realistic browser headers, triggering that specific block-block-content... block.

Here are actionable solutions to get past this:

1. Add Proper Browser Headers to Bypass Basic Detection

First, switch to requests (it’s more intuitive than urllib) and add headers that mimic a real browser. This can stop the site from immediately blocking you.

Example code:

import requests
from bs4 import BeautifulSoup

url = 'https://meetinglibrary.asco.org/browse-meetings/2019%20Gastrointestinal%20Cancers%20Symposium'

# Mimic a Chrome browser request
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://meetinglibrary.asco.org/'
}

response = requests.get(url, headers=headers)
page_soup = BeautifulSoup(response.text, "html.parser")

# Check if the block is gone
print(page_soup.find('div', class_='block-block-content'))

2. Use a JavaScript Rendering Tool (Critical for Angular)

Since the content is rendered client-side, static HTTP requests won’t capture the div.inner elements. You need a tool that loads the page like a real browser, runs JavaScript, and fetches the fully rendered HTML.

Option A: Selenium (Most Widely Used)

Install Selenium and a browser driver (e.g., ChromeDriver) first:

pip install selenium

Example code:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup

url = 'https://meetinglibrary.asco.org/browse-meetings/2019%20Gastrointestinal%20Cancers%20Symposium'

# Configure headless Chrome (no visible window)
chrome_options = Options()
chrome_options.add_argument('--headless=new')
chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')

driver = webdriver.Chrome(options=chrome_options)
driver.get(url)

# Wait for Angular to finish rendering content (adjust timeout if needed)
driver.implicitly_wait(10)

# Grab the fully rendered HTML
page_html = driver.page_source
page_soup = BeautifulSoup(page_html, "html.parser")

# Now you should access the div.inner containers
containers = page_soup.findAll("div", {"class": "inner"})
for container in containers:
    print(container.text.strip())

driver.quit()

Option B: Playwright (Modern Alternative)

Playwright handles dynamic content smoothly and is easier to set up:

pip install playwright
playwright install chrome

Example code:

from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup

url = 'https://meetinglibrary.asco.org/browse-meetings/2019%20Gastrointestinal%20Cancers%20Symposium'

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(user_agent='Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
    # Wait until network is quiet (signals Angular finished rendering)
    page.goto(url, wait_until='networkidle')
    
    page_html = page.content()
    page_soup = BeautifulSoup(page_html, "html.parser")
    
    containers = page_soup.findAll("div", {"class": "inner"})
    for container in containers:
        print(container.text.strip())
    
    browser.close()

3. Avoid Getting Banned Long-Term

  • Add delays: Insert time.sleep(2) between requests to mimic human pacing.
  • Rotate proxies: If you’re scraping multiple pages, use proxy services to avoid IP blocks.
  • Check robots.txt: Always review https://meetinglibrary.asco.org/robots.txt to ensure you’re allowed to scrape the content you’re targeting.

Final Note

ASCO may have terms of service restricting scraping, so make sure you’re complying with their rules before proceeding. If possible, look for an official API—this is a more reliable and ethical way to access meeting data if available.

内容的提问来源于stack exchange,提问作者Debayan Dutta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:06:58