使用Python+BeautifulSoup导出表格到txt仅保存最后一行问题求助
问题原因
listToStr变量在for循环内部被重复赋值,每次遍历新表格行时,上一行的处理结果会被直接覆盖- 文件写入操作放在for循环外部,仅会执行一次,最终写入的是循环结束后最后一次被赋值的
listToStr内容
解决方案
提供两种常用修复方案,可根据场景选择:
方案1:循环内逐行写入(适合数据量较大的场景)
将文件写入逻辑移到循环内部,提前打开文件避免频繁IO损耗:
html = await page.content() soup = bs.BeautifulSoup(html,'lxml') table = soup.find('table', attrs={'class':'a-bordered a-horizontal-stripes mt-table'}) table_rows = table.find_all('tr') with open('readme.txt', 'w', encoding='utf-8') as f: for tr in table_rows: td = tr.find_all('td', attrs={'data-column' : 'purchase_order'}) row = [tr.text for tr in td] listToStr = ' '.join([str(elem) for elem in row]) print(type(listToStr)) # 逐行写入,添加换行符区分不同行数据 f.write(listToStr + '\n')
方案2:先汇总所有行再统一写入(适合数据量较小的场景)
先用列表存储所有行的处理结果,最后一次性写入文件:
html = await page.content() soup = bs.BeautifulSoup(html,'lxml') table = soup.find('table', attrs={'class':'a-bordered a-horizontal-stripes mt-table'}) table_rows = table.find_all('tr') all_rows = [] for tr in table_rows: td = tr.find_all('td', attrs={'data-column' : 'purchase_order'}) row = [tr.text for tr in td] listToStr = ' '.join([str(elem) for elem in row]) print(type(listToStr)) all_rows.append(listToStr) with open('readme.txt', 'w', encoding='utf-8') as f: f.write('\n'.join(all_rows))
注:添加
encoding='utf-8'参数可以避免特殊字符、中文乱码问题。
内容的提问来源于stack exchange,提问作者Zafar Abbas
相关产品推荐
相关产品推荐

