如何在Pandas中从含多圆点的文件名提取文件扩展名?
To get the file extension (the part after the last dot) from your FileName column, your initial split-by-dots approach was on the right track—but you need to access the last element of each split list per row, not the last row of the Series. Here are a few robust ways to solve this:
Method 1: Use str.split() + str[-1]
This is the simplest approach. When you split each filename by dots, you get a list of parts for every row. Pandas' str accessor lets you grab the last element of each list directly:
import pandas as pd # Sample DataFrame X = pd.DataFrame({'FileName': ['a.b.c.d.txt', 'j.k.l.exe']}) # Extract file extension X['FileType'] = X.FileName.str.split('.').str[-1]
Result:
| FileName | FileType |
|---|---|
| a.b.c.d.txt | txt |
| j.k.l.exe | exe |
Note: This returns an empty string for filenames ending with a dot (e.g., file.) and the full filename if there are no dots (e.g., readme).
Method 2: Use Regular Expressions with str.extract()
For stricter control (like handling cases without valid extensions), use a regex to capture text after the last dot. This returns NaN for filenames with no extension:
X['FileType'] = X.FileName.str.extract(r'\.([^.]+)$', expand=False)
The regex breakdown:
\.: Matches a literal dot([^.]+): Captures one or more non-dot characters (the extension)$: Ensures we target the last dot in the string
Note: This returns NaN for filenames like readme (no dots) or .bashrc (hidden file with no extension).
Method 3: Use os.path.splitext() (Native Path Handling)
If you need to handle full file paths or rely on Python's built-in path utilities, use apply() with os.path.splitext:
import os X['FileType'] = X.FileName.apply(lambda x: os.path.splitext(x)[1].lstrip('.'))
os.path.splitext() returns a tuple (root, ext) where ext starts with a dot (e.g., .txt). We use lstrip('.') to remove the leading dot and get just the extension name.
Note: This works seamlessly even if FileName includes full paths (e.g., /home/user/docs/report.pdf returns pdf).
Why Your Initial Code Failed
When you tried X.FileName.str.split(pat='.')[-1], you were accessing the last row of the Series of lists, not the last element of each list in every row. Similarly, pop(-1) is a list method that can't be applied directly to a Series—you need Pandas' vectorized operations (like str[-1]) or apply() to handle each row individually.
内容的提问来源于stack exchange,提问作者Shridhar R Kulkarni

