You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬虫函数无法向TXT文件写入内容求助

为什么你的Python爬虫生成空TXT文件?

Hey there! Let's break down why your script is leaving you with empty text files, even though you swear it worked before. The core issue boils down to loop logic that's preventing your script from finishing (and thus closing the files properly)—plus a few other small tweaks that'll make your code more reliable.

The Root Cause

Looking at your code, the line that updates to_test_links is tucked inside the try block:

to_test_links = (list(set(links) ^ set(tested_links)))

That means if processing a link throws an error (and you hit the except block), this line never runs. Your while loop keeps checking len(to_test_links) > 0—but since to_test_links never gets updated, the loop runs forever.

Since your files are only closed at the very end of the function (all_urls.close()), if the loop never exits, the file buffers never get flushed to disk. That's why you see empty files—your writes are stuck in memory, not saved to the drive.

Fixes to Get Your Script Working Again

Here's how to adjust your code step by step:

No matter if a link succeeds or fails, you need to update the list of links to test. Move that logic outside the try so it runs every time.

2. Use with Statements for File Handling

with automatically closes files when the block ends, even if an error occurs. This avoids data loss from unclosed files and eliminates the need for manual close() calls.

3. Narrow Down Exception Handling

Catching every possible exception with except: hides bugs. Instead, catch specific network-related errors (like URLError or HTTPError) so you can debug issues without masking other problems.

4. Move Imports Outside the Function

Importing modules inside a function runs every time the function is called—better to put them at the top of your script for efficiency.

Revised Code

from urllib.request import urlopen, Request
from urllib.error import URLError, HTTPError
from bs4 import BeautifulSoup

def get_subpages(url):
    # Gets all subpage links from a website that start with the given url
    links = [url]
    tested_links = []
    to_test_links = links.copy()  # Use copy to avoid modifying the original list
    
    # Use with statements to auto-close files
    with open('all_urls.txt', 'w') as all_urls, open('problematic_pages.txt', 'w') as problematic_pages:
        while len(to_test_links) > 0:
            # Create a copy to iterate over, since we'll modify to_test_links
            current_test_links = to_test_links.copy()
            to_test_links = []
            
            for link in current_test_links:
                print('Testing link:', link)
                tested_links.append(link)
                
                try:
                    # Write to file immediately
                    all_urls.write(f"{link}\n")
                    
                    req = Request(link)
                    html_page = urlopen(req)
                    soup = BeautifulSoup(html_page, features="html.parser")
                    
                    templinks = []
                    for sublink in soup.findAll('a'):
                        href = sublink.get('href')
                        if href is not None:
                            templinks.append(href)
                    
                    # Add valid new sublinks to our main list
                    for templink in templinks:
                        if templink.startswith(url) and templink not in links:
                            links.append(templink)
                    
                    # Collect untested links for the next loop
                    new_links = set(links) - set(tested_links)
                    to_test_links.extend(new_links)
                    
                except (URLError, HTTPError) as e:
                    problematic_pages.write(f"{link}\n")
                    print(f"ERROR with {link}: {str(e)}. Check manually if needed.")
                except Exception as e:
                    # Catch unexpected errors without breaking the script
                    problematic_pages.write(f"{link}\n")
                    print(f"Unexpected error with {link}: {str(e)}")
    
    return links

Additional Notes

  • I changed templink.find(url) == 0 to templink.startswith(url)—it's more readable and does the exact same check.
  • Using set(links) - set(tested_links) is clearer than symmetric difference (^), since tested_links is always a subset of links.
  • Copying to_test_links before iterating avoids weird behavior from modifying the list while looping.

Give this revised code a try—it should properly write links to your text files and exit the loop when all links are tested.

内容的提问来源于stack exchange,提问作者curies

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 13:13:15