You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理CSV文本遇TypeError:float对象不可迭代求助

问题:清理文本标点时触发TypeError: 'float' object is not iterable错误

导入.csv文件后操作正常,定义移除标点的函数后,调用data["verified_reviews"].apply(punc)时触发错误,相关代码及错误栈如下:

# import the data using read_csv
data = pd.read_csv("text_data.csv")
data

# Let's define a function to remove punctuations

def punc(message):
    no_punc = [char for char in message if char not in string.punctuation]
    join_punc = "".join(no_punc)

    return join_punc

# Let's remove punctuations from our dataset 
data["new_verified_reviews"] = data["verified_reviews"].apply(punc)

错误栈:

**TypeError                                 Traceback (most recent call last)**
Cell In[62], line 2
      1 # Let's remove punctuations from our dataset 
----> 2 data["new_verified_reviews"] = data["verified_reviews"].apply(punc)

File ~\AppData\Local\anaconda3\envs\myenv\lib\site-packages\pandas\core\series.py:4753, in Series.apply(self, func, convert_dtype, args, by_row, **kwargs)
   4625 def apply(
   4626     self,
   4627     func: AggFuncType,
   (...)
   4632     **kwargs,
   4633 ) -> DataFrame | Series:
   4634     """
   4635     Invoke function on values of Series.
   4636 
   (...)
   4751     dtype: float64
   4752     """
-> 4753     return SeriesApply(
   4754         self,
   4755         func,
   4756         convert_dtype=convert_dtype,
   4757         by_row=by_row,
   4758         args=args,
   4759         kwargs=kwargs,
   4760     ).apply()

File ~\AppData\Local\anaconda3\envs\myenv\lib\site-packages\pandas\core\apply.py:1207, in SeriesApply.apply(self)
   1204     return self.apply_compat()
   1206 # self.func is Callable
-> 1207 return self.apply_standard()

File ~\AppData\Local\anaconda3\envs\myenv\lib\site-packages\pandas\core\apply.py:1287, in SeriesApply.apply_standard(self)
   1281 # row-wise access
   1282 # apply doesn't have a `na_action` keyword and for backward compat reasons
   1283 # we need to give `na_action="ignore"` for categorical data.
   1284 # TODO: remove the `na_action="ignore"` when that default has been changed in
   1285 #  Categorical (GH51645).
   1286 action = "ignore" if isinstance(obj.dtype, CategoricalDtype) else None
-> 1287 mapped = obj._map_values(
   1288     mapper=curried, na_action=action, convert=self.convert_dtype
   1289 )
   1291 if len(mapped) and isinstance(mapped[0], ABCSeries):
   1292     # GH#43986 Need to do list(mapped) in order to get treated as nested
   1293     #  See also GH#25959 regarding EA support
   1294     return obj._constructor_expanddim(list(mapped), index=obj.index)

File ~\AppData\Local\anaconda3\envs\myenv\lib\site-packages\pandas\core\base.py:921, in IndexOpsMixin._map_values(self, mapper, na_action, convert)
    918 if isinstance(arr, ExtensionArray):
    919     return arr.map(mapper, na_action=na_action)
-> 921 return algorithms.map_array(arr, mapper, na_action=na_action, convert=convert)

File ~\AppData\Local\anaconda3\envs\myenv\lib\site-packages\pandas\core\algorithms.py:1814, in map_array(arr, mapper, na_action, convert)
   1812 values = arr.astype(object, copy=False)
   1813 if na_action is None:
-> 1814     return lib.map_infer(values, mapper, convert=convert)
   1815 else:
   1816     return lib.map_infer_mask(
   1817         values, mapper, mask=isna(values).view(np.uint8), convert=convert
   1818     )

File lib.pyx:2917, in pandas._libs.lib.map_infer()

Cell In[59], line 5, in punc(words)
      4 def punc(words):
----> 5     no_punc = [char for char in words if char not in string.punctuation]
      6     join_punc = "".join(no_punc)
      8     return join_punc

TypeError: 'float' object is not iterable

原因分析

错误根源是verified_reviews列中存在空值(NaN),NaN在pandas中以float类型存储。当apply遍历到这些NaN值时,函数punc尝试遍历float类型的NaN,而float不可迭代,因此触发TypeError。

解决方案

方法1:修改函数,处理非字符串输入

在函数中先判断输入是否为字符串,若不是(比如NaN),直接返回空字符串或原值:

import string
import pandas as pd

def punc(message):
    # 先判断是否为字符串类型
    if not isinstance(message, str):
        return ""  # 或返回message,根据需求选择
    no_punc = [char for char in message if char not in string.punctuation]
    join_punc = "".join(no_punc)
    return join_punc

data["new_verified_reviews"] = data["verified_reviews"].apply(punc)

方法2:提前清理空值

在处理前先删除或填充verified_reviews列的空值:

# 删除含空值的行
data = data.dropna(subset=["verified_reviews"])
# 或用空字符串填充空值
data["verified_reviews"] = data["verified_reviews"].fillna("")

# 再执行标点移除
data["new_verified_reviews"] = data["verified_reviews"].apply(punc)

方法3:使用pandas内置str方法(更高效)

pandas的字符串方法会自动处理NaN,无需额外判断,代码更简洁高效:

import string
import re

# 创建标点正则表达式并替换
punc_pattern = "[" + re.escape(string.punctuation) + "]"
data["new_verified_reviews"] = data["verified_reviews"].str.replace(punc_pattern, "", regex=True)

或者直接用str.translate:

translator = str.maketrans("", "", string.punctuation)
data["new_verified_reviews"] = data["verified_reviews"].str.translate(translator)

内容的提问来源于stack exchange,提问作者Luis Enrique Orozco Villanueva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 23:30:28