编程新手求教:如何编写代码抓取网站内容并保存为TXT文件?
Hey there! As someone new to programming, Python is hands down your best bet for web scraping—it’s super readable, has tons of beginner-friendly resources, and there are great libraries that do most of the heavy lifting for you. Let’s break this down step by step:
First, you’ll need two key libraries:
requests: Sends requests to the website to fetch its contentBeautifulSoup: Parses the messy HTML code into something easy to work with
Open your terminal/command prompt and run this command to install them:
pip install requests beautifulsoup4
Here’s a complete, commented example that scrapes a website, extracts key info, and saves it to a TXT file. I’ll use example.com as a placeholder—just swap it with your target site!
# Import the tools we need import requests from bs4 import BeautifulSoup # Replace this with the URL you want to scrape target_url = "https://example.com" try: # Send a request to the website to get its content response = requests.get(target_url) # Make sure the request worked (status code 200 = success) response.raise_for_status() # Turn the raw HTML into a parseable object soup = BeautifulSoup(response.text, "html.parser") # Customize this part to extract what you need! # Example 1: Grab the page title page_title = soup.title.string # Example 2: Grab all paragraph text from the page all_paragraphs = [p.get_text(strip=True) for p in soup.find_all("p")] # Format the content we want to save saved_content = f"Scraped from: {target_url}\n\n" saved_content += f"Page Title: {page_title}\n\n" saved_content += "Paragraphs:\n" + "\n---\n".join(all_paragraphs) # Save to a TXT file (utf-8 ensures special characters work) with open("scraped_results.txt", "w", encoding="utf-8") as file: file.write(saved_content) print("Success! Check scraped_results.txt for your content.") # Catch and print any errors (like broken links or blocked requests) except Exception as e: print(f"Oops, something went wrong: {e}")
To scrape specific content (like product prices, news headlines, etc.), you’ll need to inspect the website’s HTML:
- Right-click on the element you want to scrape and select "Inspect"
- Look for unique identifiers like
classoridattributes - Update the
soup.find()orsoup.find_all()lines. For example, if headlines have a classnews-headline, use:headlines = [h.get_text(strip=True) for h in soup.find_all("h2", class_="news-headline")]
- Check if scraping is allowed: Add
/robots.txtto the end of the website’s URL (e.g.,https://example.com/robots.txt) to see if they prohibit scraping certain pages. - Avoid getting blocked: Some sites flag automated requests. Add a "user agent" header to mimic a real browser:
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"} response = requests.get(target_url, headers=headers) - Don’t spam requests: Add a small delay between requests with
time.sleep(2)(you’ll need toimport timefirst) to avoid overwhelming the website’s server.
内容的提问来源于stack exchange,提问作者RandomPerson

