如何用Python获取谷歌搜索页面IP并抓取搜索结果页面数据
Hey there! Great job diving into raw sockets to fetch web pages— that's such a solid way to learn how HTTP works under the hood. Let's break down your question step by step.
Can you do this with connect() and recv()?
Absolutely! But you need to tweak your HTTP request to include the search keyword, and fix a few small details to make it play nice with Google's servers. Here's what you need to know:
- Google's search results live at the path
/search, with aqparameter for your keyword (like/search?q=python+programming). - HTTP/1.1 requires a
Hostheader in your request— skip this, and Google will reject your request outright. - You’ll need to read the full response (not just 5000 bytes) since search results are way longer than that.
- Heads up: Google almost always redirects HTTP requests to HTTPS, so you might need to use SSL sockets if you want the actual search page (more on that below).
Modified Socket Code
Here's an updated version of your code that fetches search results and saves them to a file:
import socket import sys try: mysock = socket.socket(socket.AF_INET, socket.SOCK_STREAM) except socket.error: print("Failed to create socket.") sys.exit() try: # Google's HTTP port is 80, but they'll redirect to HTTPS host = (socket.gethostbyname("www.google.com"), 80) except socket.gaierror: print("Failed to get host") sys.exit() mysock.connect(host) # Replace "python+programming" with your desired search keyword search_keyword = "python+programming" # Construct the full HTTP GET request with required headers message = f"GET /search?q={search_keyword} HTTP/1.1\r\nHost: www.google.com\r\nConnection: close\r\n\r\n".encode() try: mysock.sendall(message) except socket.error: print("Failed to send") sys.exit() # Receive the full response (loop until no more data comes in) full_data = b"" while True: data = mysock.recv(5000) if not data: break full_data += data mysock.close() # Write the raw HTTP response to a file with open("google_search_results_raw.txt", "wb") as f: f.write(full_data) print("Search results saved to google_search_results_raw.txt")
Quick Notes on This Code:
- The
Connection: closeheader tells Google to close the socket after sending the response, so our loop knows when to stop reading. - You’ll see the response includes a
301 Moved Permanentlyredirect to HTTPS. To get the actual search page, you’d need to use an SSL socket (socket.SSLContext) and connect to port 443— that adds a bit more complexity.
Better Alternative: Use the requests Library
Raw sockets are awesome for learning, but for real-world web tasks, using a library like requests will save you tons of headache. It handles redirects, HTTPS, headers, and response parsing automatically.
First, install requests if you haven’t already:
pip install requests
Simple Requests Code for Google Search
import requests search_keyword = "python+programming" url = f"https://www.google.com/search?q={search_keyword}" # Add a user-agent header to avoid being blocked by Google's anti-scrape measures headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36" } response = requests.get(url, headers=headers) # Check if the request was successful if response.status_code == 200: with open("google_search_results.html", "w", encoding="utf-8") as f: f.write(response.text) print("Search results saved to google_search_results.html") else: print(f"Request failed with status code: {response.status_code}")
Critical Heads-Up
Google has strict anti-scraping policies. Don’t send too many requests too quickly, and always include a valid User-Agent header to mimic a real browser. For large-scale projects, consider using Google’s official Search API instead to avoid getting your IP blocked.
内容的提问来源于stack exchange,提问作者williah

