Python爬虫爬取ASCO会议页面遇内容拦截问题求助
Hey, I’ve run into similar anti-scraping blocks with Angular-powered sites before, so let’s break down what’s happening and how to fix it:
- ASCO’s site uses Angular (client-side rendering), meaning the
div.innercontent you’re targeting isn’t present in the initial raw HTML response—it gets rendered by JavaScript after the page loads. Your currenturllibcode only fetches the static base HTML, which doesn’t include the meeting data you want. - The site is also flagging your request as a bot because generic
urllibrequests lack realistic browser headers, triggering that specificblock-block-content...block.
Here are actionable solutions to get past this:
1. Add Proper Browser Headers to Bypass Basic Detection
First, switch to requests (it’s more intuitive than urllib) and add headers that mimic a real browser. This can stop the site from immediately blocking you.
Example code:
import requests from bs4 import BeautifulSoup url = 'https://meetinglibrary.asco.org/browse-meetings/2019%20Gastrointestinal%20Cancers%20Symposium' # Mimic a Chrome browser request headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://meetinglibrary.asco.org/' } response = requests.get(url, headers=headers) page_soup = BeautifulSoup(response.text, "html.parser") # Check if the block is gone print(page_soup.find('div', class_='block-block-content'))
2. Use a JavaScript Rendering Tool (Critical for Angular)
Since the content is rendered client-side, static HTTP requests won’t capture the div.inner elements. You need a tool that loads the page like a real browser, runs JavaScript, and fetches the fully rendered HTML.
Option A: Selenium (Most Widely Used)
Install Selenium and a browser driver (e.g., ChromeDriver) first:
pip install selenium
Example code:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup url = 'https://meetinglibrary.asco.org/browse-meetings/2019%20Gastrointestinal%20Cancers%20Symposium' # Configure headless Chrome (no visible window) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') driver = webdriver.Chrome(options=chrome_options) driver.get(url) # Wait for Angular to finish rendering content (adjust timeout if needed) driver.implicitly_wait(10) # Grab the fully rendered HTML page_html = driver.page_source page_soup = BeautifulSoup(page_html, "html.parser") # Now you should access the div.inner containers containers = page_soup.findAll("div", {"class": "inner"}) for container in containers: print(container.text.strip()) driver.quit()
Option B: Playwright (Modern Alternative)
Playwright handles dynamic content smoothly and is easier to set up:
pip install playwright playwright install chrome
Example code:
from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup url = 'https://meetinglibrary.asco.org/browse-meetings/2019%20Gastrointestinal%20Cancers%20Symposium' with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page(user_agent='Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') # Wait until network is quiet (signals Angular finished rendering) page.goto(url, wait_until='networkidle') page_html = page.content() page_soup = BeautifulSoup(page_html, "html.parser") containers = page_soup.findAll("div", {"class": "inner"}) for container in containers: print(container.text.strip()) browser.close()
3. Avoid Getting Banned Long-Term
- Add delays: Insert
time.sleep(2)between requests to mimic human pacing. - Rotate proxies: If you’re scraping multiple pages, use proxy services to avoid IP blocks.
- Check
robots.txt: Always reviewhttps://meetinglibrary.asco.org/robots.txtto ensure you’re allowed to scrape the content you’re targeting.
Final Note
ASCO may have terms of service restricting scraping, so make sure you’re complying with their rules before proceeding. If possible, look for an official API—this is a more reliable and ethical way to access meeting data if available.
内容的提问来源于stack exchange,提问作者Debayan Dutta

