You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

编程新手求教:如何编写代码抓取网站内容并保存为TXT文件?

Hey there! As someone new to programming, Python is hands down your best bet for web scraping—it’s super readable, has tons of beginner-friendly resources, and there are great libraries that do most of the heavy lifting for you. Let’s break this down step by step:

Step 1: Install Required Libraries

First, you’ll need two key libraries:

  • requests: Sends requests to the website to fetch its content
  • BeautifulSoup: Parses the messy HTML code into something easy to work with

Open your terminal/command prompt and run this command to install them:

pip install requests beautifulsoup4
Step 2: Basic Scraping & Saving to TXT

Here’s a complete, commented example that scrapes a website, extracts key info, and saves it to a TXT file. I’ll use example.com as a placeholder—just swap it with your target site!

# Import the tools we need
import requests
from bs4 import BeautifulSoup

# Replace this with the URL you want to scrape
target_url = "https://example.com"

try:
    # Send a request to the website to get its content
    response = requests.get(target_url)
    # Make sure the request worked (status code 200 = success)
    response.raise_for_status()

    # Turn the raw HTML into a parseable object
    soup = BeautifulSoup(response.text, "html.parser")

    # Customize this part to extract what you need!
    # Example 1: Grab the page title
    page_title = soup.title.string
    # Example 2: Grab all paragraph text from the page
    all_paragraphs = [p.get_text(strip=True) for p in soup.find_all("p")]

    # Format the content we want to save
    saved_content = f"Scraped from: {target_url}\n\n"
    saved_content += f"Page Title: {page_title}\n\n"
    saved_content += "Paragraphs:\n" + "\n---\n".join(all_paragraphs)

    # Save to a TXT file (utf-8 ensures special characters work)
    with open("scraped_results.txt", "w", encoding="utf-8") as file:
        file.write(saved_content)

    print("Success! Check scraped_results.txt for your content.")

# Catch and print any errors (like broken links or blocked requests)
except Exception as e:
    print(f"Oops, something went wrong: {e}")
Step 3: Customize for Your Target Site

To scrape specific content (like product prices, news headlines, etc.), you’ll need to inspect the website’s HTML:

  1. Right-click on the element you want to scrape and select "Inspect"
  2. Look for unique identifiers like class or id attributes
  3. Update the soup.find() or soup.find_all() lines. For example, if headlines have a class news-headline, use:
    headlines = [h.get_text(strip=True) for h in soup.find_all("h2", class_="news-headline")]
    
Important Tips for Beginners
  • Check if scraping is allowed: Add /robots.txt to the end of the website’s URL (e.g., https://example.com/robots.txt) to see if they prohibit scraping certain pages.
  • Avoid getting blocked: Some sites flag automated requests. Add a "user agent" header to mimic a real browser:
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"}
    response = requests.get(target_url, headers=headers)
    
  • Don’t spam requests: Add a small delay between requests with time.sleep(2) (you’ll need to import time first) to avoid overwhelming the website’s server.

内容的提问来源于stack exchange,提问作者RandomPerson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 07:57:45