You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup提取成语页面的指定内容?

解决方案

问题分析

原代码的核心问题:

  • 详情页请求未携带headers,易触发反爬导致请求失败
  • 未从详情页提取目标数据(标题、含义、示例、起源)
  • 未将提取的数据存入对应字母的分类列表中

修正后的代码

import requests
from bs4 import BeautifulSoup
import json

headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36'}
base_url = "https://www.englishclub.com/ref/Idioms/"

# 生成A到W的字母列表,替代手动输入
letter_list = [chr(ord('A') + i) for i in range(23)]
idioms_dict = {letter: [] for letter in letter_list}

for letter in letter_list:
    category_url = f"{base_url}{letter}/"
    try:
        # 请求字母分类页面
        category_response = requests.get(category_url, headers=headers)
        category_response.raise_for_status()  # 捕获HTTP请求错误
        category_soup = BeautifulSoup(category_response.text, "html.parser")
        idiom_links = category_soup.select('.linktitle a')
        
        for link in idiom_links:
            idiom_url = link['href']
            # 补全相对路径为完整URL
            if not idiom_url.startswith('http'):
                idiom_url = f"https://www.englishclub.com{idiom_url}"
            
            try:
                # 请求成语详情页
                idiom_response = requests.get(idiom_url, headers=headers)
                idiom_response.raise_for_status()
                idiom_soup = BeautifulSoup(idiom_response.text, "html.parser")
                main_content = idiom_soup.find('main')
                
                # 提取成语标题
                title = main_content.find('h1').get_text(strip=True) if main_content.find('h1') else None
                
                # 提取含义
                meaning = None
                meaning_h2 = main_content.find('h2', string='Meaning')
                if meaning_h2:
                    meaning_p = meaning_h2.find_next_sibling('p')
                    meaning = meaning_p.get_text(strip=True) if meaning_p else None
                
                # 提取示例
                examples = []
                examples_h2 = main_content.find('h2', string='Examples')
                if examples_h2:
                    examples_ul = examples_h2.find_next_sibling('ul')
                    if examples_ul:
                        examples = [li.get_text(strip=True) for li in examples_ul.find_all('li')]
                
                # 提取起源(可能包含多个段落)
                origin = None
                origin_h2 = main_content.find('h2', string='Origin')
                if origin_h2:
                    origin_paragraphs = []
                    next_element = origin_h2.find_next_sibling()
                    # 收集所有后续p标签,直到遇到下一个h2
                    while next_element and next_element.name != 'h2':
                        if next_element.name == 'p':
                            origin_paragraphs.append(next_element.get_text(strip=True))
                        next_element = next_element.find_next_sibling()
                    origin = '\n'.join(origin_paragraphs) if origin_paragraphs else None
                
                # 将数据存入对应字母的列表
                if title:
                    idioms_dict[letter].append({
                        'title': title,
                        'meaning': meaning,
                        'examples': examples,
                        'origin': origin
                    })
                    
            except Exception as e:
                print(f"处理成语链接 {idiom_url} 出错: {e}")
                continue
                
    except Exception as e:
        print(f"处理分类页面 {category_url} 出错: {e}")
        continue

# 保存为JSON文件
with open('idioms.json', 'w', encoding='utf-8') as f:
    json.dump(idioms_dict, f, ensure_ascii=False, indent=4)

print("爬取完成,数据已保存到idioms.json")

关键修改说明

  • 统一请求头:所有请求携带headers,避免被网站拦截
  • 完整URL处理:自动补全相对路径为完整URL,避免无效请求
  • 精准定位数据:通过h2标签的文本内容定位对应模块,适配页面结构
  • 异常处理:添加try-except捕获请求和解析错误,保证程序稳定运行
  • 结构化存储:将每个成语的信息整理为字典,按字母分类存入最终JSON

内容的提问来源于stack exchange,提问作者NewPartizal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 08:05:25