如何在Python中用正则匹配2/5字符及高效提取Pandas列中发动机缸数
问题1:用Python正则匹配长度为2或5个字符的内容
直接用正则表达式匹配恰好2个字符或恰好5个字符的内容,核心规则是限定字符串首尾的长度:
- 正则表达式:
^.{2}$|^.{5}$^和$分别匹配字符串的开头和结尾,确保是完整匹配而非部分匹配.{2}匹配任意2个字符,.{5}匹配任意5个字符|表示逻辑“或”
示例代码:
import re test_strings = ["ab", "abcde", "abc", "abcd", "abcdef"] pattern = r'^.{2}$|^.{5}$' matches = [s for s in test_strings if re.match(pattern, s)] print(matches) # 输出: ['ab', 'abcde']
如果需要匹配字符串中的子串(而非完整字符串),去掉首尾的^和$即可:
pattern = r'.{2}|.{5}' test_str = "abcdefghij" matches = re.findall(pattern, test_str) print(matches) # 输出: ['ab', 'cdefg', 'hi']
问题2:批量提取Pandas DataFrame Engine列的发动机缸数
不用逐个编写格式规则,直接提取字符串中的数字并筛选4-12的有效缸数即可,因为所有格式(I4、4 Cyl、V4、4 Cylinder等)的核心都是数字:
步骤1:提取数字并转换类型
用str.extract提取字符串中的第一个数字,转成数值类型:
import pandas as pd # 示例数据 engine_data = { "Engine": ["I4 Turbo Engine", "6 Cylinder", "V8", "12-Cyl Performance", "3 Cyl Eco", "V10"], "Count": [150, 200, 100, 50, 80, 70] } df = pd.DataFrame(engine_data) # 提取数字 df['Cylinders'] = df['Engine'].str.extract(r'(\d+)').astype(float)
步骤2:筛选有效缸数(4-12)
只保留4到12之间的数值,其余设为NaN:
df['Cylinders'] = df['Cylinders'].where(df['Cylinders'].between(4, 12))
最终结果
处理后的DataFrame:
| Engine | Count | Cylinders |
|---|---|---|
| I4 Turbo Engine | 150 | 4.0 |
| 6 Cylinder | 200 | 6.0 |
| V8 | 100 | 8.0 |
| 12-Cyl Performance | 50 | 12.0 |
| 3 Cyl Eco | 80 | NaN |
| V10 | 70 | 10.0 |
如果想一步完成提取和筛选,也可以直接用正则匹配4-12的数字:
df['Cylinders'] = df['Engine'].str.extract(r'(4|5|6|7|8|9|10|11|12)').astype(float)
内容的提问来源于stack exchange,提问作者Scott
相关产品推荐
相关产品推荐

