使用BeautifulSoup爬取网页表格无法获取某列全部数据
问题根因
soup.find_all('table', class_="wikitable")返回的是页面中所有class为wikitable的整张表格的集合,不是表格内的行。你要爬取的频率列表本身就是1个独立的wikitable,所以你的for循环实际只执行了1次。- 循环里调用
t.find('span', class_="Hans")只会返回当前(整张表格)范围内第一个匹配的span元素,也就是第一行的“一”,自然拿不到后续所有词条。 - 额外bug:你写csv表头时把
simplified和pinyin写在了同一个字符串里,导出后会变成单列,不是预期的两列结构。
修正方案
- 直接用
find()定位到唯一的目标wikitable即可,不需要用find_all()拿表格集合 - 提取表格内所有
<tr>行标签,跳过第一行表头后逐行遍历 - 从每一行中分别提取简体字、拼音对应的元素内容,逐行写入csv
修正后的可运行代码:
from bs4 import BeautifulSoup import requests import csv wiki_url = "https://en.wiktionary.org/wiki/Appendix:Mandarin_Frequency_lists/1-1000" response = requests.get(wiki_url) response.encoding = 'utf-8' soup = BeautifulSoup(response.text, 'html.parser') # 定位目标频率表 table = soup.find('table', class_="wikitable") with open('chinesewords.csv', 'w', encoding='utf8', newline='') as f: writer = csv.writer(f) # 写入正确的两列表头 writer.writerow(["simplified", "pinyin"]) # 跳过表头行,遍历所有数据行 for row in table.find_all('tr')[1:]: simplified = row.find('span', class_="Hans").text.strip() pinyin = row.find('span', class_="pinyin").text.strip() writer.writerow([simplified, pinyin]) print(simplified, pinyin)
运行后会按顺序打印1000个词条的简体和拼音,csv文件也会按两列格式正确存储所有数据。
内容的提问来源于stack exchange,提问作者jeffshoe
相关产品推荐
相关产品推荐

