You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫抓取希腊语词形页面时表格内容提取错误的解决方法

修正希腊语词形分析爬虫的内容提取问题

问题根源

你遇到的核心问题是搜索范围未限定——直接调用全局find_all时,会把页面所有符合class="code"的元素全部抓取,而非仅在当前lemmacontainer表格内查找,导致后续表格的结果叠加了前面表格的内容,出现数量不符的情况。

通用修正方案

遍历每个lemmacontainer表格时,必须将搜索范围严格限定在当前表格内部,而非全局搜索。以下是调整后的代码逻辑:

示例修正代码

import requests
from bs4 import BeautifulSoup

# 替换为你的目标页面URL
target_url = "需要抓取的页面地址"
response = requests.get(target_url)
soup = BeautifulSoup(response.text, 'html.parser')

# 先定位所有目标表格
lemmatables = soup.find_all('table', class_='lemmacontainer')

# 逐个表格处理,限定搜索范围
for table_index, table in enumerate(lemmatables, start=1):
    # 关键:仅在当前表格内查找class为code的元素
    code_items = table.find_all('span', class_='code')
    code_texts = [item.get_text(strip=True) for item in code_items]
    
    print(f"第{table_index}个表格的code内容:")
    print(code_texts)
    print(f"提取数量:{len(code_texts)}\n")

额外优化建议

  • 若页面存在动态加载内容(如AJAX异步加载、滚动加载),requests无法获取完整DOM,需改用selenium或playwright模拟浏览器渲染。
  • 检查页面源码确认lemmacontainer和code的class拼写完全一致,部分页面可能存在大小写、空格或嵌套class的情况。
  • 添加基础异常处理,避免因页面结构变动导致脚本崩溃,比如判断code_items是否为空后再执行后续逻辑。

内容的提问来源于stack exchange,提问作者user21978357

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 08:17:09