如何用Python正则findall结合月份列表批量匹配文本日期
Solution to Batch Match Dates with Month Names/Abbreviations Using
findall Got it, let's sort this out. The issue with your initial approach is that you're trying to loop through each month and call findall separately, which won't work as a single line (and would be inefficient anyway). Instead, we can convert your months list into a regex alternation group (using | to separate options) so we can match all month variants in one go.
Step-by-Step Explanation
- Convert the months list to a regex pattern: We'll join all month names/abbreviations with
|, escaping each one to avoid any unexpected regex behavior (though your months don't have special characters, this is a safe practice). - Build the full date regex: Attach the pattern for day and year to the month group, making sure to capture each part (month, day, year) as separate groups.
- Run
findallonce: Apply the combined regex to your DataFrame column to extract all matching dates in one line.
Working Code Example
import pandas as pd import re # Your month list months = ['January', 'February', 'March', 'April', 'May', 'June', 'July', 'August', 'September', 'October','November', 'December', 'Jan', 'Feb', 'Mar', 'Apr', 'Jun', 'Jul', 'Aug', 'Sep', 'Oct', 'Nov', 'Dec'] # Sample DataFrame df = pd.DataFrame({ 'text': ['last day of the championship July 28, 1983 Mar 11, 1990 record of the first division April 27, 1982 record of played matches'] }) # Build the regex pattern month_pattern = '|'.join(re.escape(month) for month in months) date_regex = rf'({month_pattern})\s(\d?\d),\s(\d{{4}})' # Extract all matches df['extracted_dates'] = df['text'].str.findall(date_regex) # View the result print(df['extracted_dates'].iloc[0])
Expected Output
[('July', '28', '1983'), ('Mar', '11', '1990'), ('April', '27', '1982')]
Why This Works
- The
month_patternbecomes a string likeJanuary|February|...|Dec, so regex will match any of these month variants. - The full
date_regexcaptures three groups: the month (full or abbreviated), the day (1 or 2 digits), and the 4-digit year. str.findall()returns all non-overlapping matches as tuples, exactly the format you're looking for.
If you want to condense this into a single line (as you asked), you can combine the pattern building and findall call:
df['extracted_dates'] = df['text'].str.findall(r'(' + '|'.join(re.escape(m) for m in months) + r')\s(\d?\d),\s(\d{4})')
内容的提问来源于stack exchange,提问作者Joe
相关产品推荐
相关产品推荐

