如何为含空格的字符串添加前后缀?(Pandas实现报错求助)
仅对DataFrame中path列包含单个空格的行,为该列字符串添加^前缀和$后缀。
错误尝试代码
df[df.path.str.contains(" ", na=False)] = df['^' + df.path.astype('str') + '$']
报错信息
KeyError Traceback (most recent call last)
Cell In[180], line 22
17 df['path'] = df.path.str.replace(" ", "$|^")
18 #print(df[df.path.str.contains("$|^", na=False)].shape[0])
19
20 #multipath = df[df.path.str.contains(" ", na=False)]
21 #multipath['path'] = multipath['path'].str.replace(' ','$|^')
---> 22 df[df.path.str.contains(" ", na=False)] = df['^' + df.path.astype('str') + '$']
23 #multipath['path'] =
24 #print(multipath.shape[0])
27 print("original shape:", df.shape[0])File ~/Library/Python/3.9/lib/python/site-packages/pandas/core/frame.py:3902, in DataFrame.getitem(self, key)
3900 if is_iterator(key):
3901 key = list(key)
-> 3902 indexer = self.columns._get_indexer_strict(key, "columns")[1]
3904 # take() does not accept boolean indexers
3905 if getattr(indexer, "dtype", None) == bool:File ~/Library/Python/3.9/lib/python/site-packages/pandas/core/indexes/base.py:6114, in Index._get_indexer_strict(self, key, axis_name)
6111 else:
6112 keyarr, indexer, new_indexer = self._reindex_non_unique(keyarr)
-> 6114 self._raise_if_missing(keyarr, indexer, axis_name)
6116 keyarr = self.take(indexer)
6117 if isinstance(key, Index):
6118 # GH 42790 - Preserve name from an IndexFile ~/Library/Python/3.9/lib/python/site-packages/pandas/core/indexes/base.py:6175, in Index._raise_if_missing(self, key, indexer, axis_name)
6173 if use_interval_msg:
6174 key = list(key)
-> 6175 raise KeyError(f"None of [{key}] are in the [{axis_name}]")
6177 not_found = list(ensure_index(key)[missing_mask.nonzero()[0]].unique())
6178 raise KeyError(f"{not_found} not in index")KeyError: "None of [Index(['^nan$',\n '^/content/icebreakers/en_us/products/ice-cubes-wintergreen-1-5-oz-tins-8-ct-box.html$',\n '^/content/corporate_SSF/en_us/careers/retirees/contact-us.html$',\n '^/content/franchise/en_us/products.html$',\n '^/content/franchise/en_us/our-brands/good-and-plenty/good-and-plenty-1-8-oz.html$',\n '^/content/dam/chocolateworld/en_us/documents/times-square-reopening-menu.pdf$',\n '^/content/franchise/en_us/products/hersheys-milk-chocolate-snack-size-bag.html$',\n '^/content/jolly-rancher/en_us/products/orignal-hard-candy-60-oz.html$',\n '^/content/kitkat/en_us/products/white-bar.html$',\n '^/content/dam/chocolateworld/en_us/documents/birthday-party-packages.pdf$',\n ...\n '^/kitchens/en_us/recipes/grilled-peanut-butter-chocolate-and-jelly-sandwich.html$',\n '^/kitchens/en_us/recipes/apple-wheels.html$',\n '^/whoppers/en_us/products/strawberry-milkshake.html$',\n '^/kitchens/en_us/blogs/chocolate-covered-valentines-day.html$',\n '/en_us/our-brands/bubble-yum/bubble-yum-cotton-candy-5-piece.html$|/en_us/our-brands/bubble-yum/bubble-yum-sugarless-original-10-piece.html$',\n '^/en_us/ad-cookie-policy.html$',\n '^/hersheysolutions/en_us/employee_register.html$',\n '/reeses/en_us/products/$|/en_us/our-brands/milk-duds/milk-duds-snack-size-chewy-caramels-9-3-oz.html$|/reeses/en_us/products/reeses-peanut-butter-eggs-6-pack.html$|/reeses/en_us/products/reeses-chocolate-lovers-cups.html$',\n '^nan$', '^/content/corporate_SSF/en_us/investors.html$'],\n dtype='object', length=4586)] are in the [columns]"
错误原因
df['^' + df.path.astype('str') + '$']是在尝试用拼接后的字符串作为列名取值,但这些字符串并非DataFrame的列名,直接触发KeyError。- 原代码试图替换整行数据,而非仅修改
path列,逻辑错误。 contains(" ")会匹配所有含空格的行,无法精准筛选仅含单个空格的行。
正确实现方式
方法1:用loc定位+正则匹配
# 正则匹配仅含单个空格的行:开头到空格无其他空格,空格到结尾无其他空格 mask = df['path'].str.contains(r'^[^ ]* [^ ]*$', na=False) # 仅修改符合条件的path列值 df.loc[mask, 'path'] = '^' + df.loc[mask, 'path'].astype(str) + '$'
方法2:apply结合空格计数
def process_path(s): if isinstance(s, str) and s.count(' ') == 1: return f'^{s}$' return s df['path'] = df['path'].apply(process_path)
方法3:np.where批量处理
import numpy as np df['path'] = np.where( df['path'].str.count(' ') == 1, '^' + df['path'].astype(str) + '$', df['path'] )
内容的提问来源于stack exchange,提问作者Ramy

