Python网页爬虫函数无法向TXT文件写入内容求助
Hey there! Let's break down why your script is leaving you with empty text files, even though you swear it worked before. The core issue boils down to loop logic that's preventing your script from finishing (and thus closing the files properly)—plus a few other small tweaks that'll make your code more reliable.
The Root Cause
Looking at your code, the line that updates to_test_links is tucked inside the try block:
to_test_links = (list(set(links) ^ set(tested_links)))
That means if processing a link throws an error (and you hit the except block), this line never runs. Your while loop keeps checking len(to_test_links) > 0—but since to_test_links never gets updated, the loop runs forever.
Since your files are only closed at the very end of the function (all_urls.close()), if the loop never exits, the file buffers never get flushed to disk. That's why you see empty files—your writes are stuck in memory, not saved to the drive.
Fixes to Get Your Script Working Again
Here's how to adjust your code step by step:
1. Move to_test_links Update Outside the Try/Except Block
No matter if a link succeeds or fails, you need to update the list of links to test. Move that logic outside the try so it runs every time.
2. Use with Statements for File Handling
with automatically closes files when the block ends, even if an error occurs. This avoids data loss from unclosed files and eliminates the need for manual close() calls.
3. Narrow Down Exception Handling
Catching every possible exception with except: hides bugs. Instead, catch specific network-related errors (like URLError or HTTPError) so you can debug issues without masking other problems.
4. Move Imports Outside the Function
Importing modules inside a function runs every time the function is called—better to put them at the top of your script for efficiency.
Revised Code
from urllib.request import urlopen, Request from urllib.error import URLError, HTTPError from bs4 import BeautifulSoup def get_subpages(url): # Gets all subpage links from a website that start with the given url links = [url] tested_links = [] to_test_links = links.copy() # Use copy to avoid modifying the original list # Use with statements to auto-close files with open('all_urls.txt', 'w') as all_urls, open('problematic_pages.txt', 'w') as problematic_pages: while len(to_test_links) > 0: # Create a copy to iterate over, since we'll modify to_test_links current_test_links = to_test_links.copy() to_test_links = [] for link in current_test_links: print('Testing link:', link) tested_links.append(link) try: # Write to file immediately all_urls.write(f"{link}\n") req = Request(link) html_page = urlopen(req) soup = BeautifulSoup(html_page, features="html.parser") templinks = [] for sublink in soup.findAll('a'): href = sublink.get('href') if href is not None: templinks.append(href) # Add valid new sublinks to our main list for templink in templinks: if templink.startswith(url) and templink not in links: links.append(templink) # Collect untested links for the next loop new_links = set(links) - set(tested_links) to_test_links.extend(new_links) except (URLError, HTTPError) as e: problematic_pages.write(f"{link}\n") print(f"ERROR with {link}: {str(e)}. Check manually if needed.") except Exception as e: # Catch unexpected errors without breaking the script problematic_pages.write(f"{link}\n") print(f"Unexpected error with {link}: {str(e)}") return links
Additional Notes
- I changed
templink.find(url) == 0totemplink.startswith(url)—it's more readable and does the exact same check. - Using
set(links) - set(tested_links)is clearer than symmetric difference (^), sincetested_linksis always a subset oflinks. - Copying
to_test_linksbefore iterating avoids weird behavior from modifying the list while looping.
Give this revised code a try—it should properly write links to your text files and exit the loop when all links are tested.
内容的提问来源于stack exchange,提问作者curies

