使用Pandas.read_html因Emoji触发ParserError:文档为空的解决求助
解决Pandas.read_html()因Emoji触发的lxml解析错误
在用Pandas.read_html()爬取网页表格时,HTML源码中的Emoji会触发如下解析错误:
lxml.etree.ParserError: Document is empty
尝试过直接读取HTML源码字符串、提取<table>标签解析均失败,爬取流程使用Selenium+BeautifulSoup,搭配html5lib和lxml作为解析依赖,因爬取页面体量较大(约100页),无法手动移除Emoji,需自动处理。
可复现示例
import pandas as pd table_tag = """<table> <tr> <th>Company</th> <th>Contact</th> <th>Country</th> </tr> <tr> <td>Notfall Software 🚑</td> <td>Mario Müller</td> <td>Germany</td> </tr> <tr> <td>Centro comercial Pélican</td> <td>Francisco Villa 😅</td> <td>Mexico</td> </tr> </table>""" tables = pd.read_html(table_tag, encoding='utf-8') # 此处触发错误 df = tables[0]
手动移除Emoji后可正常生成DataFrame:
>>> df Company Contact Country 0 Notfall Software Mario Müller Germany 1 Centro comercial Pélican Francisco Villa Mexico
解决方案
方法1:解析前用正则自动移除Emoji
通过正则匹配所有Emoji字符,在传入pd.read_html()前清理HTML字符串:
import pandas as pd import re # 匹配全范围Emoji的正则表达式 emoji_pattern = re.compile( "[" u"\U0001F600-\U0001F64F" # 表情符号 u"\U0001F300-\U0001F5FF" # 符号&图标 u"\U0001F680-\U0001F6FF" # 交通&地图符号 u"\U0001F1E0-\U0001F1FF" # 国旗 u"\U00002500-\U00002BEF" # 线条、几何图形 u"\U00002702-\U000027B0" u"\U000024C2-\U0001F251" u"\U0001f926-\U0001f937" u"\U00010000-\U0010ffff" u"\u2640-\u2642" u"\u2600-\u2B55" u"\u200d" u"\u23cf" u"\u23e9" u"\u231a" u"\ufe0f" # 变体选择符 u"\u3030" "]+", flags=re.UNICODE ) # 清理HTML字符串中的Emoji cleaned_table_tag = emoji_pattern.sub(r'', table_tag) # 正常解析 tables = pd.read_html(cleaned_table_tag, encoding='utf-8') df = tables[0] print(df)
方法2:指定html5lib作为解析器
html5lib对特殊字符(如Emoji)的兼容性优于lxml,直接指定解析器即可跳过错误:
tables = pd.read_html(table_tag, encoding='utf-8', flavor='html5lib') df = tables[0]
注:html5lib解析速度比lxml慢,但100页的体量完全可接受。
方法3:通过BeautifulSoup精准清理文本节点
若需保留HTML结构仅清理文本中的Emoji,可遍历BeautifulSoup节点处理:
from bs4 import BeautifulSoup import pandas as pd import re emoji_pattern = re.compile( "[" u"\U0001F600-\U0001F64F" u"\U0001F300-\U0001F5FF" u"\U0001F680-\U0001F6FF" u"\U0001F1E0-\U0001F1FF" u"\U00002500-\U00002BEF" u"\U00002702-\U000027B0" u"\U000024C2-\U0001F251" u"\U0001f926-\U0001f937" u"\U00010000-\U0010ffff" u"\u2640-\u2642" u"\u2600-\u2B55" u"\u200d" u"\u23cf" u"\u23e9" u"\u231a" u"\ufe0f" u"\u3030" "]+", flags=re.UNICODE ) # 用BeautifulSoup加载HTML soup = BeautifulSoup(table_tag, 'html.parser') # 遍历所有文本节点,移除Emoji for text_node in soup.find_all(text=True): cleaned_text = emoji_pattern.sub(r'', text_node) text_node.replace_with(cleaned_text) # 转为字符串后解析 cleaned_table_tag = str(soup) tables = pd.read_html(cleaned_table_tag, encoding='utf-8') df = tables[0] print(df)
内容的提问来源于stack exchange,提问作者Domenico Spidy Tamburro
相关产品推荐
相关产品推荐

