Python实现谷歌搜索链接批量抓取及新闻搜索方法
Hey there! Let's walk through how to handle both of your Google scraping tasks with Beautiful Soup.
Google uses the start parameter to manage pagination—each page displays roughly 10 results, so the pattern is:
- Page 1:
start=0(default, no need to explicitly include it) - Page 2:
start=10 - Page 3:
start=20 - And so on, incrementing by 10 for each subsequent page.
Here's a practical code snippet that loops through multiple pages:
import requests from bs4 import BeautifulSoup from urllib.parse import urlparse, parse_qs import time def scrape_google_pages(query, num_pages): base_url = "http://www.google.ca/search" # Set a real user-agent to avoid immediate blocking headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } all_links = [] for page in range(num_pages): start = page * 10 params = { "q": query, "start": start } # Add a delay to be respectful of Google's servers time.sleep(2) response = requests.get(base_url, params=params, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # Extract search result links (adjust selectors if Google updates its HTML) result_containers = soup.find_all("div", class_="g") for container in result_containers: link_tag = container.find("a") if link_tag and "href" in link_tag.attrs: raw_link = link_tag["href"] # Parse the actual URL from Google's redirect wrapper if raw_link.startswith("/url?q="): parsed_url = urlparse(raw_link) actual_link = parse_qs(parsed_url.query)["q"][0] all_links.append(actual_link) return all_links # Example: Scrape 3 pages of results for "python web scraping tips" links = scrape_google_pages("python web scraping tips", 3) for link in links: print(link)
Key reminders:
- Always include a valid
User-Agent—without it, Google will block your requests instantly. - Don't skip the
time.sleep()call—firing too many requests too fast will get your IP flagged. - Google's HTML structure can change over time, so you might need to tweak the selector (like
class_="g") if the code stops working.
To switch exclusively to Google News, just add the tbm=nws parameter to your search URL. This tells Google to return only news content, and you can still use the start parameter for pagination here.
Here's the adjusted code for Google News scraping:
def scrape_google_news(query, num_pages): base_url = "http://www.google.ca/search" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } all_news_links = [] for page in range(num_pages): start = page * 10 params = { "q": query, "tbm": "nws", # This flag enables news-only search "start": start } time.sleep(2) response = requests.get(base_url, params=params, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # News containers might have minor structural differences—adjust if needed news_containers = soup.find_all("div", class_="g") for container in news_containers: link_tag = container.find("a") if link_tag and "href" in link_tag.attrs: raw_link = link_tag["href"] if raw_link.startswith("/url?q="): parsed_url = urlparse(raw_link) actual_link = parse_qs(parsed_url.query)["q"][0] all_news_links.append(actual_link) return all_news_links # Example: Scrape 2 pages of news for "AI ethics developments" news_links = scrape_google_news("AI ethics developments", 2) for link in news_links: print(link)
A quick heads-up: If you're planning to scrape at scale, Google's anti-scraping systems are tough to bypass long-term. For production projects, the official Google Custom Search API is a far more reliable (and TOS-compliant) option. But for small, personal projects, the above methods will work as long as you're mindful of request rates.
内容的提问来源于stack exchange,提问作者Matthew Macfarlane

