You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas.read_html因Emoji触发ParserError:文档为空的解决求助

解决Pandas.read_html()因Emoji触发的lxml解析错误

在用Pandas.read_html()爬取网页表格时,HTML源码中的Emoji会触发如下解析错误:

lxml.etree.ParserError: Document is empty

尝试过直接读取HTML源码字符串、提取<table>标签解析均失败,爬取流程使用Selenium+BeautifulSoup,搭配html5lib和lxml作为解析依赖,因爬取页面体量较大(约100页),无法手动移除Emoji,需自动处理。

可复现示例

import pandas as pd

table_tag = """<table>
  <tr>
    <th>Company</th>
    <th>Contact</th>
    <th>Country</th>
  </tr>
  <tr>
    <td>Notfall Software 🚑</td>
    <td>Mario Müller</td>
    <td>Germany</td>
  </tr>
  <tr>
    <td>Centro comercial Pélican</td>
    <td>Francisco Villa 😅</td>
    <td>Mexico</td>
  </tr>
</table>"""

tables = pd.read_html(table_tag, encoding='utf-8') # 此处触发错误
df = tables[0]

手动移除Emoji后可正常生成DataFrame:

>>> df
                    Company          Contact  Country
0          Notfall Software     Mario Müller  Germany
1  Centro comercial Pélican  Francisco Villa   Mexico

解决方案

方法1:解析前用正则自动移除Emoji

通过正则匹配所有Emoji字符,在传入pd.read_html()前清理HTML字符串:

import pandas as pd
import re

# 匹配全范围Emoji的正则表达式
emoji_pattern = re.compile(
    "["
    u"\U0001F600-\U0001F64F"  # 表情符号
    u"\U0001F300-\U0001F5FF"  # 符号&图标
    u"\U0001F680-\U0001F6FF"  # 交通&地图符号
    u"\U0001F1E0-\U0001F1FF"  # 国旗
    u"\U00002500-\U00002BEF"  # 线条、几何图形
    u"\U00002702-\U000027B0"
    u"\U000024C2-\U0001F251"
    u"\U0001f926-\U0001f937"
    u"\U00010000-\U0010ffff"
    u"\u2640-\u2642" 
    u"\u2600-\u2B55"
    u"\u200d"
    u"\u23cf"
    u"\u23e9"
    u"\u231a"
    u"\ufe0f"  # 变体选择符
    u"\u3030"
                  "]+", 
    flags=re.UNICODE
)

# 清理HTML字符串中的Emoji
cleaned_table_tag = emoji_pattern.sub(r'', table_tag)
# 正常解析
tables = pd.read_html(cleaned_table_tag, encoding='utf-8')
df = tables[0]
print(df)

方法2:指定html5lib作为解析器

html5lib对特殊字符(如Emoji)的兼容性优于lxml,直接指定解析器即可跳过错误:

tables = pd.read_html(table_tag, encoding='utf-8', flavor='html5lib')
df = tables[0]

注:html5lib解析速度比lxml慢,但100页的体量完全可接受。

方法3:通过BeautifulSoup精准清理文本节点

若需保留HTML结构仅清理文本中的Emoji,可遍历BeautifulSoup节点处理:

from bs4 import BeautifulSoup
import pandas as pd
import re

emoji_pattern = re.compile(
    "["
    u"\U0001F600-\U0001F64F"
    u"\U0001F300-\U0001F5FF"
    u"\U0001F680-\U0001F6FF"
    u"\U0001F1E0-\U0001F1FF"
    u"\U00002500-\U00002BEF"
    u"\U00002702-\U000027B0"
    u"\U000024C2-\U0001F251"
    u"\U0001f926-\U0001f937"
    u"\U00010000-\U0010ffff"
    u"\u2640-\u2642"
    u"\u2600-\u2B55"
    u"\u200d"
    u"\u23cf"
    u"\u23e9"
    u"\u231a"
    u"\ufe0f"
    u"\u3030"
                  "]+", 
    flags=re.UNICODE
)

# 用BeautifulSoup加载HTML
soup = BeautifulSoup(table_tag, 'html.parser')
# 遍历所有文本节点,移除Emoji
for text_node in soup.find_all(text=True):
    cleaned_text = emoji_pattern.sub(r'', text_node)
    text_node.replace_with(cleaned_text)

# 转为字符串后解析
cleaned_table_tag = str(soup)
tables = pd.read_html(cleaned_table_tag, encoding='utf-8')
df = tables[0]
print(df)

内容的提问来源于stack exchange,提问作者Domenico Spidy Tamburro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 15:17:31