You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取求助:提取The Hacker News文章标题与链接并邮件推送

解决方案:过滤冗余标签,仅保留文章标题与链接

问题分析

你的代码目前存在两个核心问题导致邮件出现冗余HTML标签:

  1. 直接抓取页面所有<a>标签,包含了很多非文章的导航链接、广告链接,且每个标签本身带有多余的嵌套元素;
  2. Redis存储的是完整<a>标签的字符串,邮件中直接拼接这些原始标签,自然会引入冗余内容。

修改步骤与代码优化

1. 精准定位文章标题与链接

The Hacker News的文章结构是每个文章包裹在div.post块里,标题位于h2.home-title下的<a>标签,直接定位这些元素就能拿到纯净的文章标题和链接:

def parse(self):
    soup = BeautifulSoup(self.markup, 'html.parser')
    # 定位所有文章块
    posts = soup.find_all('div', class_='post')
    self.saved_links = []
    for post in posts:
        # 提取文章标题链接
        title_tag = post.find('h2', class_='home-title').find('a')
        if title_tag:
            title = title_tag.get_text(strip=True)
            url = title_tag['href']
            # 检查标题是否包含关键词(大小写不敏感)
            for keyword in self.keywords:
                if keyword.lower() in title.lower():
                    self.saved_links.append({'title': title, 'url': url})
    print(self.saved_links)

2. 修改Redis存储逻辑

不再存储完整HTML标签,而是存储标题与链接的对应关系,方便后续构建干净的邮件内容:

def store(self):
    r = redis.Redis(host='localhost', port=6379, db=0)
    for item in self.saved_links:
        # 用标题作为键,链接作为值存储
        r.set(item['title'], item['url'])

3. 重构邮件内容生成逻辑

从Redis取出标题和链接后,手动构建干净的<a>标签,避免冗余内容:

def email(self):
    r = redis.Redis(host='localhost', port=6379, db=0)
    # 获取所有标题和对应链接
    links = []
    for title in r.keys():
        title_str = title.decode('utf-8')
        url = r.get(title).decode('utf-8')
        # 构建干净的HTML链接
        links.append(f'<a href="{url}" target="_blank">{title_str}</a>')

    # 邮件部分代码保持原有逻辑,仅修改HTML拼接
    import smtplib
    from email.mime.multipart import MIMEMultipart
    from email.mime.text import MIMEText

    fromEmail = "your-email@gmail.com"  # 替换为实际发件邮箱
    toEmail = "recipient-email@gmail.com"  # 替换为实际收件邮箱

    msg = MIMEMultipart('alternative')
    msg['Subject'] = "今日感兴趣的The Hacker News文章"
    msg['From'] = fromEmail
    msg['To'] = toEmail

    # 构建干净的邮件HTML内容
    html = f"""
        <h4>共找到 {len(links)} 篇你可能感兴趣的文章:</h4>
        {'<br/><br/>'.join(links)}
    """

    mime = MIMEText(html, 'html')
    msg.attach(mime)

    try:
        mail = smtplib.SMTP('smtp.gmail.com', 587)
        mail.ehlo()
        mail.starttls()
        mail.login(fromEmail, bot_email_pw)
        mail.sendmail(fromEmail, toEmail, msg.as_string())
        mail.quit()
        print('邮件发送成功!')
    except Exception as exc:
        print(f'发送失败:{exc}')

    r.flushdb()

4. 完整优化后代码

from bs4 import BeautifulSoup
import redis
from password import bot_email_pw
import requests

class Scraper:
    def __init__(self, keywords):
        self.markup = requests.get('https://thehackernews.com/').text
        self.keywords = keywords

    def parse(self):
        soup = BeautifulSoup(self.markup, 'html.parser')
        posts = soup.find_all('div', class_='post')
        self.saved_links = []
        for post in posts:
            title_tag = post.find('h2', class_='home-title').find('a')
            if title_tag:
                title = title_tag.get_text(strip=True)
                url = title_tag['href']
                for keyword in self.keywords:
                    if keyword.lower() in title.lower():
                        self.saved_links.append({'title': title, 'url': url})
        print(self.saved_links)

    def store(self):
        r = redis.Redis(host='localhost', port=6379, db=0)
        for item in self.saved_links:
            r.set(item['title'], item['url'])

    def email(self):
        r = redis.Redis(host='localhost', port=6379, db=0)
        links = []
        for title in r.keys():
            title_str = title.decode('utf-8')
            url = r.get(title).decode('utf-8')
            links.append(f'<a href="{url}" target="_blank">{title_str}</a>')

        import smtplib
        from email.mime.multipart import MIMEMultipart
        from email.mime.text import MIMEText

        fromEmail = "your-email@gmail.com"
        toEmail = "recipient-email@gmail.com"

        msg = MIMEMultipart('alternative')
        msg['Subject'] = "今日感兴趣的The Hacker News文章"
        msg['From'] = fromEmail
        msg['To'] = toEmail

        html = f"""
            <h4>共找到 {len(links)} 篇你可能感兴趣的文章:</h4>
            {'<br/><br/>'.join(links)}
        """

        mime = MIMEText(html, 'html')
        msg.attach(mime)

        try:
            mail = smtplib.SMTP('smtp.gmail.com', 587)
            mail.ehlo()
            mail.starttls()
            mail.login(fromEmail, bot_email_pw)
            mail.sendmail(fromEmail, toEmail, msg.as_string())
            mail.quit()
            print('邮件发送成功!')
        except Exception as exc:
            print(f'发送失败:{exc}')

        r.flushdb()

# 运行示例
s = Scraper(['malware'])
s.parse()
s.store()
s.email()

关键优化点

  • 精准爬取:通过文章块div.post定位,只抓取真正的文章标题链接,过滤无效链接;
  • 纯净存储:Redis仅存储标题和链接文本,不存HTML标签;
  • 干净邮件:手动构建<a>标签,确保邮件内容只有标题和跳转链接,无冗余元素。

内容的提问来源于stack exchange,提问作者Dawid Herman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 01:06:20