Python爬虫问题:如何去除冗余字符并按行写入爬取的邮箱
Python爬虫问题修复方案
问题梳理
- 爬取的邮箱无法按每个链接对应内容写入新行
- 输出文件中出现
['\n这类不属于邮箱的冗余字符
问题根源与修复步骤
1. 无法写入新行的原因及修复
原代码打开文件时使用默认只读模式,且未添加换行符,导致内容直接拼接或无法正常写入。修复方式:
- 使用追加模式(
"a")打开文件,确保新内容追加到文件末尾而非覆盖原有内容 - 每个邮箱写入后添加换行符
"\n",保证每行一个独立邮箱
2. 冗余字符的原因及修复
原代码直接将邮箱列表转成字符串写入,导致列表的格式符号([、'、,)被写入文件。修复方式:
- 遍历邮箱列表,逐个写入每个邮箱
- 优化正则匹配规则,清理结果中的多余空格和换行符
修改后的完整代码
import os import random import re import requests from bs4 import BeautifulSoup def scrapeEmails(): global reqs, _lock, success, fails, rps, rpm with open(os.path.join("proxies.txt"), "r") as f: proxies = f.read().splitlines() with open(os.path.join("links_toscrape.txt"), "r") as f: channelLinks = f.read().splitlines() rndChannelLinks = random.choice(channelLinks) URL = rndChannelLinks + "/about" proxy = random.choice(proxies) proxies = {"https": "http://"+proxy} try: soup = BeautifulSoup(requests.get(URL, proxies=proxies).text, "html.parser") _description = soup.find("meta", property="og:description") _content = _description["content"] if _description else "No meta title given" if "@" in _content.lower(): # 优化正则,精准匹配合法邮箱格式 __email = re.findall(r"([\w.+-]{1,63}@[\w.-]{1,63})", _content) # 清理邮箱中的空格、换行符 cleanEmail = [x.strip().replace("\n", "") for x in __email] print("Email: ", cleanEmail) # 追加模式打开文件,逐个写入邮箱并换行 with open("scraped_emails.txt", "a", encoding="utf-8") as f: for email in cleanEmail: f.write(f"{email}\n") else: print(f"Email of YouTube channel {URL} not found.") except Exception as e: print(f"Error processing {URL}: {str(e)}")
关键修改说明
- 文件操作:改用
"a"追加模式,指定utf-8编码避免乱码,with语句自动管理文件关闭,无需手动调用f.close() - 内容写入:遍历邮箱列表,每行写入一个邮箱,彻底解决新行问题
- 正则优化:调整正则表达式,支持更多合法邮箱字符(如
+、-),同时用strip()清理前后空格 - 异常处理:增加
try-except块捕获请求或解析错误,避免程序意外崩溃
内容的提问来源于stack exchange,提问作者Jane1990
相关产品推荐
相关产品推荐

