You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R语言dplyr包处理MLB比赛linescore列:提取各局得分并转换为数值列的技术问询

Parsing MLB Linescore Data with dplyr & stringr

Hey there! Let's work through your three goals for parsing the MLB linescore data, all using dplyr (plus stringr for regex handling, which plays perfectly with dplyr). Let's start with your existing code and fix/expand it step by step.

1. Extract First Inning Score & Fix Parenthesis Removal

Your initial code to pull the first inning score was spot-on, but the error with removing parentheses comes because you didn’t wrap str_remove_all inside the mutate() call. Dplyr doesn’t recognize the new inng1 column until the mutate() block finishes, so we need to handle both steps in one go:

library(dplyr)
library(stringr)

# Your original dataset
gamedata <- structure(list(team = c("NYM", "NYM", "BOS", "NYM", "BOS"), linescore = c("010000000", "(10)1140006x", "002200010", "00000(11)01x", "311200"), ondate = structure(c(18475, 18476, 18487, 18489, 18494), class = "Date")), class = "data.frame", row.names = c(NA, -5L))

# Extract first inning and clean parentheses in a single mutate step
gamedata_first_inning <- gamedata %>%
  select(ondate, team, linescore) %>%
  mutate(
    # Grab the first inning: either a single digit or a two-digit value in parentheses
    inng1_raw = str_extract(linescore, "^(\\d|\\(\\d{2}\\))"),
    # Remove parentheses and convert to numeric
    inng1 = as.numeric(str_remove_all(inng1_raw, "[()]"))
  ) %>%
  select(-inng1_raw) # Drop the raw intermediate column if you don't need it

This fixes the "object 'inng1' not found" error because we’re modifying the new column immediately within the same mutate() context.

2. Extract All Subsequent Innings into Numeric Columns

To parse every inning (including multi-digit scores and trailing x markers), we need to split the linescore into individual inning units, clean them, and reshape the data into wide format. Here’s the full workflow:

gamedata_full_innings <- gamedata %>%
  select(ondate, team, linescore) %>%
  # First, remove the trailing 'x' (marks incomplete 9th inning)
  mutate(linescore_clean = str_remove(linescore, "x$")) %>%
  # Split linescore into individual inning units:
  # Uses regex lookbehind/lookahead to split between valid score chunks
  mutate(innings = str_split(linescore_clean, "(?<=\\d|\\))(?=\\d|\\()")) %>%
  # Unnest the list column to one row per inning per game
  tidyr::unnest_longer(innings) %>%
  # Clean parentheses, assign inning numbers, convert to numeric
  mutate(
    inning_num = row_number(), # Auto-number innings (1, 2, 3...)
    runs = as.numeric(str_remove_all(innings, "[()]"))
  ) %>%
  # Pivot back to wide format with columns like inng1, inng2, etc.
  tidyr::pivot_wider(
    names_from = inning_num,
    values_from = runs,
    names_prefix = "inng"
  )

Key Details on the Regex Split

The str_split uses a positive lookbehind/lookahead regex ((?<=\\d|\\))(?=\\d|\\()) to correctly split linescores like (10)1140006 into ["(10)", "1", "1", "4", "0", "0", "0", "6"]—it splits right after a digit or closing parenthesis, and right before a digit or opening parenthesis.

Final Output Preview

Your resulting data will look like this (truncated for clarity):

ondateteamlinescoreinng1inng2inng3inng4inng5inng6inng7inng8inng9
2020-08-01NYM010000000010000000
2020-08-02NYM(10)1140006x101140006NA
2020-08-13BOS002200010002200010

All inng* columns are numeric, so you can easily calculate total runs with something like rowSums(select(., starts_with("inng")), na.rm = TRUE).

内容的提问来源于stack exchange,提问作者Metsfan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 22:57:41