You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取网页表格无法获取某列全部数据

问题根因
  • soup.find_all('table', class_="wikitable")返回的是页面中所有class为wikitable的整张表格的集合,不是表格内的行。你要爬取的频率列表本身就是1个独立的wikitable,所以你的for循环实际只执行了1次。
  • 循环里调用t.find('span', class_="Hans")只会返回当前(整张表格)范围内第一个匹配的span元素,也就是第一行的“一”,自然拿不到后续所有词条。
  • 额外bug:你写csv表头时把simplified和pinyin写在了同一个字符串里,导出后会变成单列,不是预期的两列结构。
修正方案
  • 直接用find()定位到唯一的目标wikitable即可,不需要用find_all()拿表格集合
  • 提取表格内所有<tr>行标签,跳过第一行表头后逐行遍历
  • 从每一行中分别提取简体字、拼音对应的元素内容,逐行写入csv

修正后的可运行代码:

from bs4 import BeautifulSoup
import requests
import csv

wiki_url = "https://en.wiktionary.org/wiki/Appendix:Mandarin_Frequency_lists/1-1000"
response = requests.get(wiki_url)
response.encoding = 'utf-8'
soup = BeautifulSoup(response.text, 'html.parser')

# 定位目标频率表
table = soup.find('table', class_="wikitable")

with open('chinesewords.csv', 'w', encoding='utf8', newline='') as f:
    writer = csv.writer(f)
    # 写入正确的两列表头
    writer.writerow(["simplified", "pinyin"])
    # 跳过表头行,遍历所有数据行
    for row in table.find_all('tr')[1:]:
        simplified = row.find('span', class_="Hans").text.strip()
        pinyin = row.find('span', class_="pinyin").text.strip()
        writer.writerow([simplified, pinyin])
        print(simplified, pinyin)

运行后会按顺序打印1000个词条的简体和拼音,csv文件也会按两列格式正确存储所有数据。

内容的提问来源于stack exchange,提问作者jeffshoe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 00:27:42