使用Beautiful Soup抓取HTML表格行数据失败,寻求解决方案
问题:BeautifulSoup抓取HTML表格td数据失败
我尝试用bs4抓取HTML表格,但代码无法正常运行,需要获取表格中td的行数据来写入CSV文件。目标HTML结构如下:
<table class="sc-jAaTju bVEWLO"> <thead> <tr> <td width="10%">Rank</td> <td>Trending Topic</td> <td width="30%">Tweet Volume</td> </tr> </thead> <tbody> <tr> <td>1</td> <td><a href="http:///example.com/search?q=%23One" target="_blank" without="true" rel="noopener noreferrer">#One</a></td> <td>1006.4K tweets</td> </tr> <tr> <td>2</td> <td><a href="http:///example.com/search?q=%23Two" target="_blank" without="true" rel="noopener noreferrer">#Two</a></td> <td>1028.7K tweets</td> </tr> <tr> <td>3</td> <td><a href="http:///example.com/search?q=%23Three" target="_blank" without="true" rel="noopener noreferrer">#Three</a></td> <td>Less than 10K tweets</td> </tr> </tbody> </table>
我的第一次尝试代码:
url = requests.get(f"https://www.exportdata.io/trends/italy/2020-01-01/0") soup = BeautifulSoup(url.text, "html.parser") table = soup.find_all("table", attrs={"class":"sc-jAaTju bVEWLO"})
我的第二次尝试代码:
tables = soup.find_all('table') for table in tables: td = tables.td.text.strip()
两段代码均无法正常工作,请问我忽略了什么?
问题分析与解决
两次尝试的问题点
- 第一次代码:仅完成了表格定位,没有后续提取行和单元格数据的逻辑,相当于只找到了表格,但没做数据提取。另外
find_all返回的是列表,若目标表格唯一,用find更合适。 - 第二次代码:循环中错误调用
tables.td(tables是表格列表,应使用循环变量table),且仅提取了第一个单元格,没有遍历所有行和每行内的所有td。
正确实现代码
import requests from bs4 import BeautifulSoup import csv # 发送请求并验证状态 response = requests.get("https://www.exportdata.io/trends/italy/2020-01-01/0") if response.status_code != 200: print(f"请求失败,状态码:{response.status_code}") else: soup = BeautifulSoup(response.text, "html.parser") # 定位目标表格(用class_简化属性传参) target_table = soup.find("table", class_="sc-jAaTju bVEWLO") if not target_table: print("未找到目标表格,请检查class是否正确") else: # 提取表头 headers = [header.text.strip() for header in target_table.thead.find_all("td")] # 提取tbody内的所有行 rows = target_table.tbody.find_all("tr") # 写入CSV文件 with open("trending_topics.csv", "w", newline="", encoding="utf-8") as csv_file: writer = csv.writer(csv_file) writer.writerow(headers) for row in rows: # 提取每行的所有td文本 row_content = [cell.text.strip() for cell in row.find_all("td")] writer.writerow(row_content) print("数据已成功写入CSV文件")
额外注意事项
- 如果页面数据是JS动态渲染,
requests无法获取到表格内容,需改用Selenium、Playwright等工具模拟浏览器加载页面。 - 若网站的class是动态生成(会随机变化),可换用其他定位方式,比如通过表头文本匹配表格:
# 示例:通过表头包含"Rank"和"Tweet Volume"定位表格 for table in soup.find_all("table"): thead_text = table.thead.get_text(strip=True) if table.thead else "" if "Rank" in thead_text and "Tweet Volume" in thead_text: target_table = table break
内容的提问来源于stack exchange,提问作者Kome Gognome
相关产品推荐
相关产品推荐

