Why start with StatsBomb open data

StatsBomb publishes a set of free competitions with full event data: every pass, carry, pressure and shot, with coordinates and context. It is the best sandbox in football analytics — no API key, no terms negotiation, no scraping.

The full catalogue lives in the open-data repository on GitHub.

Step 1 — Pick a competition, load the events

The statsbombpy package wraps the raw JSON:

from statsbombpy import sb

competitions = sb.competitions()
matches = sb.matches(competition_id=11, season_id=90)   # La Liga 2020/21
events = sb.events(match_id=matches.iloc[0].match_id)

If you prefer raw files — recommended for learning, because you see the schema — download events/{match_id}.json from the repository and parse it yourself with pandas.json_normalize.

Step 2 — Turn events into shots

shots = (
    events[events.type == "Shot"]
    .assign(
        x=lambda df: df.location.apply(lambda loc: loc[0]),
        y=lambda df: df.location.apply(lambda loc: loc[1]),
        goal=lambda df: df.outcome == "Goal",
    )
    [["x", "y", "goal", "player", "shot_statsbomb_xg"]]
)

print(shots.groupby("player").agg(shots=("goal", "size"), goals=("goal", "sum")))

Two things to notice: coordinates are in a 120×80 pitch, and the events already carry shot_statsbomb_xg. Use that column to validate your own model later.

Step 3 — Draw the shot map

import mplsoccer
import matplotlib.pyplot as plt

pitch = mplsoccer.Pitch(pitch_type="statsbomb", pitch_color="#0d1117")
fig, ax = pitch.draw()

sizes = shots["shot_statsbomb_xg"] * 900 + 60
colors = shots["goal"].map({True: "#00e08a", False: "#4b5563"})
pitch.scatter(shots.x, shots.y, s=sizes, c=colors, alpha=0.85, ax=ax)
ax.set_title("Shots — size = xG, colour = goal", color="white")
plt.show()

Step 4 — Make it repeatable

A pipeline that downloads, transforms and plots in one script becomes unmanageable by the third match. Split it early:

  • ingest/ — download raw JSON, store it unchanged (cache everything; the API is slow and rate-limited).
  • transform/ — parse events into tidy frames, one function per entity (shots, passes, pressures).
  • analysis/ — pure functions over the tidy frames: aggregations, rolling windows, comparisons.
  • notebooks/ — only exploration; the moment a transform is reused, it moves to transform/.

This is the same pattern every serious football-data project uses, and it is the one that survives contact with a second dataset.

Where to go next

  1. Features for xG — add pressure, body part, first-time shots; compare your AUC against shot_statsbomb_xg.
  2. Pass networks — aggregate passes by player pair and position.
  3. Pressing metrics — count pressures by zone and see which teams break structure first.

Pick one question, build the smallest pipeline that answers it, and publish what you learn. That loop — question, data, answer — is the actual job.