You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于网页爬取文件关闭、文本行截取及写入模式的技术问询

Hey there! Let's tackle your web scraping file handling and line extraction questions one by one.

Fixing File Closure & Avoiding Unwanted Appends

Your original code opens the file in append mode ("a") inside the loop, which is why you had to use append to keep all lines—using "w" there would overwrite the file every iteration, leaving only the last line. The cleanest way to handle this is to use Python's with statement, which automatically closes the file once you're done with it, no manual close() needed. Here's how to rewrite your code:

from bs4 import BeautifulSoup

# 假设soup已经是解析好的对象
with open("demo.txt", "w", encoding="utf-8") as output_file:
    for link in soup.find_all('a'):
        # 写入文本并添加换行符,避免所有内容挤在一行
        output_file.write(f"{link.text.strip()}\n")
  • Using "w" mode here is safe because we open the file once, write all lines in the loop, and the with block takes care of closing it properly. No more accidental appends on subsequent runs!
  • Adding .strip() cleans up any extra whitespace in the link text, and \n ensures each entry is on its own line.

Extracting Specific Line Ranges

To grab lines 2-5 or 94-100 from your demo.txt file, you have two solid options depending on the file size:

Option 1: Read All Lines at Once (Good for Small Files)

If your file isn't huge, you can load all lines into a list and slice it (remember Python uses 0-indexing, so adjust your line numbers accordingly):

with open("demo.txt", "r", encoding="utf-8") as input_file:
    all_lines = input_file.readlines()

# 获取2-5行(索引1到4,因为第1行对应索引0)
lines_2_to_5 = all_lines[1:5]
# 获取94-100行(索引93到99)
lines_94_to_100 = all_lines[93:100]

# 合并结果并保存到新文件
with open("extracted_lines.txt", "w", encoding="utf-8") as output_file:
    output_file.writelines(lines_2_to_5 + lines_94_to_100)

Option 2: Iterate Line by Line (Better for Large Files)

For big files that you don't want to load entirely into memory, loop through each line and check if it's in your target range:

extracted_lines = []
# 定义目标行范围为(起始行,结束行)
target_ranges = [(2, 5), (94, 100)]

with open("demo.txt", "r", encoding="utf-8") as input_file:
    # 从行号1开始枚举每一行
    for line_number, line in enumerate(input_file, start=1):
        # 检查当前行是否在目标范围内
        for start, end in target_ranges:
            if start <= line_number <= end:
                extracted_lines.append(line)
                break  # 匹配到后跳出循环,避免重复检查

# 保存提取结果
with open("extracted_lines.txt", "w", encoding="utf-8") as output_file:
    output_file.writelines(extracted_lines)

This method uses minimal memory since it only keeps the lines you need in storage.

内容的提问来源于stack exchange,提问作者Ivan Pupo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:32:20