You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式匹配行中指定词及之后的首个数字(大小写不敏感)

Got it, let's fix this regex issue for you! The problem with your original pattern (Milk+)\d is that it only looks for digits immediately following "Milk"—but we need to account for any characters in between, while still grabbing the first number that appears after the word. Here's how to do it in Python:

Step-by-Step Solution

1. The Correct Regex Pattern

We'll use a pattern that:

  • Matches "milk" (case-insensitively)
  • Skips any characters between "milk" and the first number (non-greedily, so it stops at the first digit sequence)
  • Captures both the word and the number for later use

Pattern: r'(milk).*?(\d+)'

  • (milk): Captures the target word (we'll use a flag to make this case-insensitive)
  • .*?: Non-greedy match for any characters (stops at the first occurrence of the next part)
  • (\d+): Captures one or more digits (the first number after "milk")

2. Python Code for Single Line

Here's how to apply this to your sample text:

import re

text = "Milk for human consumption may be taken only from cattle from 80 hours after the last treatment."
pattern = r'(milk).*?(\d+)'

# Search for the pattern with case-insensitive flag
match = re.search(pattern, text, re.IGNORECASE)

if match:
    # Extract the captured groups
    target_word = match.group(1)  # Will be "Milk" (matches original case)
    first_number = match.group(2)  # Will be "80"
    print(f"Word found: {target_word}, First number after it: {first_number}")
else:
    print("No match found in the line.")

3. Handling Large Files

For large files, process lines one at a time to avoid loading the entire file into memory:

import re

pattern = r'(milk).*?(\d+)'

with open("your_large_file.txt", "r") as file:
    for line_num, line in enumerate(file, 1):
        match = re.search(pattern, line, re.IGNORECASE)
        if match:
            print(f"Line {line_num}:")
            print(f"  Word: {match.group(1)}")
            print(f"  First number after word: {match.group(2)}\n")

Key Notes

  • The re.IGNORECASE flag ensures the pattern matches "milk", "Milk", "MILK", or any other case variation.
  • The non-greedy .*? is critical here—it prevents the regex from skipping past the first number to match later digits in the line.

内容的提问来源于stack exchange,提问作者darkCoffy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 19:12:49